ICLR 2026 peer review scores show a striking drop in agreement.

In a recent analysis comparing ICLR 2025 and 2026 scores, I observed striking discrepancies in the correlation between human reviewers.

3 min readMachine Learning
ICLR 2026 peer review scores show a striking drop in agreement.
Just did an analysis on ICLR 2025 vs 2026 scores and WOW [D]

The numbers are in, and they are not subtle. A Reddit user's analysis of ICLR peer review scores, drawn directly from OpenReview, shows that agreement between human reviewers has collapsed. For ICLR 2025, the correlation between two human scores was already a modest 0.41. For ICLR 2026, that figure has plummeted, with the mean within-paper human standard deviation jumping from 1.186 to 1.523. That is not a marginal shift. That is a signal that the system is breaking down in plain sight.

Let's be clear about what this means for anyone submitting to top machine learning conferences. You are not just facing a lottery where luck plays a role. You are facing a lottery where the noise has become the dominant feature. The data shows that for the average paper, two reviewers are now more likely to disagree than to align. The standard deviation within a single paper's scores has grown so large that a paper could reasonably land anywhere from a strong reject to a strong accept depending on who happens to open the file. This is not a commentary on the quality of the research. It is a commentary on the reliability of the evaluation itself.

What is striking is that the overall average score standard deviation has barely moved, from 1.253 in 2025 to 1.162 in 2026. That means the pool of scores is not getting wilder overall. The divergence is happening inside each paper. Reviewers are not collectively drifting toward harsher or more lenient standards. They are simply disagreeing with each other more often, and more severely, than they did a year ago. That is a different problem than "reviewers are too strict" or "reviewers are too lenient." It means the review process has lost its calibration.

For the researcher on the ground, the practical takeaway is uncomfortable but essential. You cannot treat a single review, or even two, as a meaningful signal about your work's value. You should expect variance, not because your paper is flawed, but because the instrument itself is noisy. The smart play is to design your submission defensively: make your contributions explicit, your experiments exhaustive, and your writing clear enough that even a distracted reviewer has less room to project their own biases onto your work. That is not gaming the system. That is surviving it.

The community should stop treating this as an isolated data point and start asking what can be done. If the peer review process cannot produce consistent scores, then the scores should carry less weight in acceptance decisions. That might mean more emphasis on author response, on reproducibility checks, or on post-publication validation. The current trajectory is not sustainable, and the data proves it. The next step is not to lament the lottery. It is to redesign the game.

From Machine Learning

Per https://paperreview.ai/tech-overview, the scores corr between 2 human is about 0.41 for ICLR 2025, but in my current project I am seeing a much lower corr for ICLR 2026. So I ran the metrics for both 2025 and 2026 and it is crazy. I used 2 metrics, one-vs-rest corr and half-half split corr. All data are fetched from OpenReview.

I do know that top conf reviews are just a lottery now for most papers, but i nenver thought it is this bad.

Read the original at Machine Learning