AI Assisted Review

AI reviews exposed: when depth meets surface in peer feedback

The review process felt misaligned with its own purpose.

4 min readMachine Learning

The review period for NeurIPS this year was always going to be a pressure test for AI-assisted peer review. But the experience shared by one participant, summarized in that Reddit thread, reveals something more uncomfortable than a few misaligned scores. It wasn't just that some reviews were shallow. It was that the tool designed to level the playing field exposed a deeper confusion about what we actually want from reviewers in the first place. When a reviewer breaks double blindness to justify a rejection by citing an LLM's specific outputs, without having engaged with the author's rebuttal, we have to ask: are we using these models to understand the work, or just to armor our own opinions?

This is the practical friction we should all care about. The author in the thread noted low clarity scores from reviewers who struggled with established notation and concepts. That is not a failure of the authors. It is a failure of the review process to set expectations for what a reviewer should know. If we are going to invite AI assistance, then the baseline for reviewer competence shifts. We can no longer accept a reviewer who flags a paper for being unclear when the issue is their own unfamiliarity with the field. This connects directly to the broader challenge of training models and humans alike, as outlined in our guide on Unlock LLM Training: A Practical Guide to Distributed Algorithms. Just as distributed training requires a shared understanding of system architecture, effective peer review requires a shared baseline of domain knowledge. The LLM should be the bridge, not the excuse.

Our take is straightforward: the double-blind condition is becoming a liability when paired with AI tools. It assumes a level playing field where none exists. The reviewer who broke anonymity to cite the LLM did not do so out of malice. They did it out of a misguided sense of justification. But the result is a process where the loudest voice is not the most informed, but the one most skilled at prompting a model to agree with them. That is not rigorous. It is performative. If we look at how LLMs navigate token space, as discussed in Exploring Paragraph Structure: How LLMs Navigate Token Space, we see that meaning is contextual and positional. The same is true in review. A reviewer's comment is only meaningful if it reflects a genuine attempt to engage with the work, not a template generated by a model.

For our readers, the takeaway is concrete. If you are an author, do not assume that low clarity scores reflect a flaw in your paper. It may be a signal that your reviewers lack the foundational context to evaluate it fairly. And if you are a reviewer, consider this: the point of AI assistance is not to give you a shortcut to a verdict. It is to help you ask better questions. The author in the thread wondered whether they should have broken anonymity to explain that the notation is standard. That would have been a mistake. The burden should not fall on the author to educate the reviewer mid-review. The burden is on the process to ensure that reviewers are equipped to handle the material, or to decline if they are not. Watch for the next cycle to see if the organizers address this gap, or if we continue to mistake generated text for genuine insight. The specific detail to watch is whether any policy emerges that holds reviewers accountable for demonstrating actual engagement with the work, not just citing the model's output.

From Machine Learning

Out of curiosity, if you were a reviewer or author, how did the review period go?

For me, it was weird, because I gave reviews with specific details (what specifically could have been better, how to fix it), but realized other reviewers gave similar superficial reviews. Even the paper which was a control for me (no LLM), I gave specific comments, but other reviewers focused on minor things.

Read the original at Machine Learning