ICLR

ICLR 2027 Resets Its Scoring Scale for Paper Reviews

ICLR's decision to compress its review scale from the familiar 1-10 down to just four scores feels like a step backward, not forward.

3 min readMachine Learning

A four-point scale for paper reviews sounds like a step toward forced clarity, but it may also be a step away from nuance. The ICLR 2027 change, compressing the review score range to 1 (Clear rejection), 2 (Weak rejection), 3 (Weak acceptance), and 4 (Clear acceptance), is a significant move that deserves more discussion than the quiet confusion on Reddit suggests. We think this reset is a deliberate attempt to reduce noise in the review process, but it raises immediate practical questions for authors and reviewers alike.

The shift mirrors broader trends in academic publishing toward tighter quality controls. The arXiv tightens submission limits to sustain quality and accessibility story shows a similar impulse: as submission volumes grow, venues are imposing constraints to preserve signal. ICLR's new scale is essentially a constraint on reviewers. By eliminating the middle ground, no 5, 6, or 7 to hedge with, the conference forces a binary-or-nearly-binary decision. That is fine for borderline papers, but what about the submission that is sound, significant, and clear, yet not earth-shattering? It now gets a 3, the same score as a paper that barely meets the bar. That compression may frustrate reviewers who want to rank contributions more precisely. Meanwhile, Navigating the Second Round of AAAI Reviews: What We Know So Far reminds us that multi-round reviews already exist to refine decisions. Perhaps ICLR's logic is that a coarse initial score plus a rebuttal phase is more honest than a false-precision 1-to-10 scale.

For the researcher submitting to ICLR 2027, the practical takeaway is clear: a score of 3 is no longer a comfortable middle. It is a weak acceptance, meaning your paper survived but did not convince. You cannot assume a 3 will lead to a poster or an oral; it might be a borderline that gets cut in the meta-review. On the flip side, a 2 is a weak rejection, not a death sentence, but a signal that the reviewers saw merit alongside flaws. Authors should treat a 2 as an invitation to engage seriously during rebuttal, not as a brush-off. The compressed scale also changes how we interpret acceptance rates. If the program chairs use the 3 and 4 scores to define the set of accepted papers, then the difference between a 3 and a 4 matters enormously. But if they treat 3 as a conditional and 4 as a firm accept, the scale introduces a new kind of uncertainty.

The real test will be whether ICLR publishes the distribution of scores after the conference. If most reviews cluster at 2 and 3, the scale has failed to create clarity. If reviews spread cleanly across all four points, it will have succeeded in simplifying a process that has long needed it. Watch for that data.

From Machine Learning

I got my three papers to review, and it seems like they have changed the review score range again this year?

–––––––––––––––––––––––––––––––––––– Based on your overall assessment of the submission, what is your recommended decision? Consider the paper’s overall soundness, significance, clarity, and contribution.

Read the original at Machine Learning