The ICML 2026 experiment with two LLM-review policies has produced an outcome that feels less like a controlled test and more like a stacked deck. Based on the anecdotal evidence shared by one author, papers reviewed under the stricter Policy A, where LLM assistance was prohibited, appear to have received harsher scores than those under the more permissive Policy B. If this pattern holds, the conference's attempt to compare review methods may have inadvertently created an uneven playing field, punishing authors who followed the rules.
What we find striking is not the possibility that LLM-assisted reviews are more lenient, but how predictable that outcome seems. The community member who raised this concern spent nearly a week crafting careful, human-only reviews, only to suspect that their batch was graded against a tougher standard. Across a small sample of about 15 Policy A papers, their score was one of the highest, yet comparable to middling results reported online from Policy B authors. The professor's suggestion that ICML will normalize scores across groups offers little comfort when the scoring itself may be systematically skewed. For authors who chose the stricter policy in good faith, this feels less like a scientific comparison and more like a penalty for caution.
The practical takeaway for anyone submitting to future conferences is uncomfortable but clear. When a review policy permits LLM assistance, it introduces a variable that goes beyond mere convenience. Polished, background-rich, and lenient reviews from Policy B may reflect the model's tendency to smooth over rough edges rather than a genuine difference in paper quality. Authors cannot control which policy their paper receives, but they should demand transparency in how scores are ultimately compared. Raw score distributions mean little if the review tools themselves shift the baseline.
We hope ICML will publish the aggregate results from the community poll that the original poster has organized. But even without that data, the warning is already written: a review policy that allows LLM assistance on one track and forbids it on another does not produce a fair comparison. It produces two different games. The conference should either standardize the policy for all papers or abandon the pretense that scores across policies are directly comparable. Authors deserve to know what standard they are being judged against, not discover it after the work is done.