The Downsides of LLM-Generated Peer Reviews [D]
Our take
The increasing integration of Large Language Models (LLMs) into academic workflows promises efficiency gains, but as a recent Reddit post highlights, [No rebuttals from neurips authors [D]] and similar experiences suggest a potential for unintended consequences. This post, detailing the pitfalls of LLM-assisted peer review, serves as a crucial cautionary tale for researchers and institutions alike. The core issue isn't simply that LLMs occasionally generate incorrect statements—it’s their tendency to produce an overwhelming volume of superficially plausible criticisms, often lacking in practical relevance or grounded in a deep understanding of the subject matter. This shift has the potential to fundamentally alter the dynamic of the peer review process, and not necessarily for the better. The article’s observation that authors are being forced to spend valuable rebuttal time addressing technically possible but practically insignificant concerns resonates deeply with anyone who’s navigated the often-arduous process of academic publication.
The article rightly points out two key problems. First, LLMs excel at identifying uncontrolled variables, generating endless lists of potential confounders without adequately assessing their actual impact on the study's conclusions. This is a logical extension of their pattern-matching abilities, but it lacks the nuanced judgment required to prioritize concerns. Secondly, LLM reviews often operate at an overly abstract level, criticizing entire research fields rather than addressing specific methodologies. For example, a critique claiming a method is “not sufficiently different from methods in Transformer” without referencing concrete papers or architectures is essentially unanswerable. This echoes concerns raised in a related piece, [Prompt, Context, Loop: The Three Engineering Layers Every RAG System Is Built On], which emphasizes the importance of grounding LLM applications in specific context and carefully controlled prompts – a lesson that clearly needs to be applied to peer review systems. The danger is that this superficiality, coupled with the sheer volume of output, will drown out genuinely insightful critiques.
The problem isn't necessarily the technology itself, but rather the uncritical adoption of LLM-generated content. The author’s warning that “copying an LLM response into a review without that judgment does not improve peer review” is particularly salient. It underscores the vital role of the human reviewer as a filter, a prioritizer, and a contextualizer. The peer review system relies on expertise and critical thinking—qualities that, at present, LLMs struggle to replicate. While LLMs can be valuable tools for assisting reviewers, they should not replace the judgment of experienced researchers. We’ve seen similar issues arise in other areas, like the recent discussion around questionable experiences reported in [Bad but typical NeurIPS experience? [D]], further highlighting the need for careful oversight and a critical approach to AI-assisted workflows. The current trend risks shifting the burden of evaluating speculative LLM output onto authors, effectively turning them into unpaid editors for AI algorithms.
Looking ahead, the challenge lies in developing strategies to harness the benefits of LLMs while mitigating their risks. This may involve implementing stricter guidelines for LLM use in peer review, emphasizing the importance of human oversight, and developing tools that help reviewers assess the relevance and validity of LLM-generated suggestions. Perhaps a future system could incorporate a “confidence score” for each LLM-generated criticism, reflecting its estimated relevance and evidentiary basis. Ultimately, the integrity of the peer review process depends on preserving the human element – the ability to discern signal from noise, to prioritize concerns, and to provide constructive feedback that genuinely advances scientific understanding. The question remains: how can we ensure that AI assists, rather than undermines, this vital function?
Having used LLMs to assist with reviews, and also having received reviews that appear to rely heavily on LLM-generated text, I have noticed two recurring problems.
1. The endless search for uncontrolled variables
LLMs are very good at identifying additional variables that were not explicitly controlled. The problem is that many of these variables have little realistic chance of changing the paper’s main conclusion.
For any experiment, it is possible to generate an almost unlimited list of potential confounders. Suppose a study finds that trees treated with fertilizer A grow better than trees treated with fertilizer B. An LLM can ask whether rainfall was perfectly controlled, whether the distribution of grass around the trees was considered, or whether wind, temperature, soil microorganisms, and countless other factors were isolated.
Each question may look logically valid in isolation. But the real issue is not whether a variable exists. The issue is whether it is sufficiently important and plausible to threaten the conclusion.
LLMs are generally poor at making this prioritization. They often convert minor residual uncertainty into what sounds like a serious methodological weakness.
This becomes especially harmful when reviewers copy such outputs directly into their reviews without independently assessing their importance. Authors are then forced to spend the rebuttal addressing an endless series of technically possible but practically insignificant concerns.
A review should not ask whether every imaginable variable has been controlled. It should ask whether the remaining uncertainty materially weakens the central claim.
2. LLM reviews are often overly abstract
Another common problem is criticism at the level of an entire research field rather than a specific prior method.
For example, an LLM may claim that a proposed method is “not sufficiently different from methods in Transformer” without identifying a concrete paper, objective, architecture, or learning relation that actually overlaps with the proposed method.
What exactly is the author expected to rebut in that situation? Every method in Transformer?
A meaningful novelty criticism should identify a specific prior method and explain precisely which components are equivalent or insufficiently differentiated. Comparing one concrete method against an entire research area is too abstract to be falsifiable or actionable.
3. LLM review is not detail
LLMs also tend to overestimate similarity between methods that share high-level terminology. Two approaches may both use architecture, concept, or attention, while differing substantially in their computational structure, training objective, assumptions, and intended use.
Because LLMs often lack a sufficiently detailed understanding of each method, they may recommend comparisons between papers that are only superficially related. The resulting review sounds comprehensive but does not demonstrate real technical understanding.
The central problem is not simply that LLM-generated reviews can contain incorrect statements. It is that they can generate an unlimited number of superficially reasonable criticisms without judging their relevance, severity, or evidentiary burden.
A strong reviewer should filter such suggestions, prioritize only the concerns that could materially affect the paper’s claims, and attach each criticism to a concrete technical basis. Copying an LLM response into a review without that judgment does not improve peer review. It merely transfers the cost of evaluating the LLM’s speculation to the authors.
[link] [comments]
Read on the original site
Open the publisher's page for the full experience