How Review Policies Shape Fairness in AI Conference Scoring

In a recent community poll regarding ICML 2026's review policies, initial findings suggest that Policy B may yield higher average scores, while Policy A reports greater reviewer confidence.

3 min readMachine Learning
How Review Policies Shape Fairness in AI Conference Scoring
[D] ICML 2026 review policy debate: 100 responses suggest Policy B may score higher, while Policy A shows higher confidence

The early data from this ICML 2026 review policy survey tells us something worth paying attention to, even if it is not a verdict. After 100 responses, Policy B papers show a higher mean score (3.43 versus 3.26), while Policy A reviews come with higher reported reviewer confidence (3.53 versus 3.35). More strikingly, 31.7% of Policy B respondents described their reviews as "especially polished," compared to just 13.6% under Policy A. These patterns do not prove that one policy caused unfair outcomes, but they do raise a practical question for anyone submitting to these conferences: how much does the review structure itself shape the scores you receive?

The inverse relationship between scores and confidence is the most telling signal. If Policy B reviews are more polished but less confident, it suggests a reviewer who may be relying on external tools to generate language they do not fully stand behind. That is not an accusation of bad faith. It is a description of how incentives work. When the review process rewards a certain kind of polished output, participants adapt. The result is a system where the same paper could receive a meaningfully different score depending on which policy bucket it lands in. For researchers, that means your outcome is partly a function of process design, not just paper quality.

The sample is self-selected and the numbers are preliminary. But the conversation is already useful. The survey creator deserves credit for turning personal frustration into a structured inquiry that the community can examine. Too often, debates about review fairness stay at the level of anecdote and accusation. Here we have a poll, a distribution, and a clear set of questions to refine. The real value is not in whether Policy B is "better" or "worse," but in whether the community now has a reason to look more carefully at how review policies interact with the tools reviewers use.

What matters next is action. If you submitted under either policy, fill out the survey and share it with colleagues who did the same. The more data the community collects, the harder it becomes to dismiss patterns like these as noise. And if future results confirm that policy design systematically influences scores, then conference organizers have a clear responsibility to adjust. That is the kind of accountability that makes peer review worth defending.

From Machine Learning

A week ago I made a thread asking whether ICML 2026’s review policy might have affected review outcomes, especially whether Policy A papers may have been judged more harshly than Policy B papers.

Original thread: https://www.reddit.com/r/MachineLearning/comments/1s387tx/d_icml_2026_policy_a_vs_policy_b_impact_on_scores/ Poll: https://docs.google.com/forms/d/e/1FAIpQLSdQilhiCx_dGLgx0tMVJ1NDX1URdJoUGIscFoPCpe6qE2Ph8w/viewform?usp=header

Read the original at Machine Learning