Peer review has always been a human exercise in judgment, but the rise of LLM-assisted reviewing is quietly testing how much of that judgment we are willing to outsource. The critique laid out here, that LLMs generate endless lists of plausible-sounding confounders and abstract field-level criticisms, is not just a technical limitation. It is a signal that we are confusing fluency with understanding. When a reviewer copies an LLM's output without filtering for relevance, they are not adding rigor; they are shifting the burden of discernment onto the authors. That is a poor trade for everyone.
The two recurring problems described deserve real attention. First, the obsession with uncontrolled variables: sure, a study on fertilizer could theoretically control for rainfall, soil microbes, or the angle of the sun, but a good reviewer knows that *theoretically possible* and *materially threatening* are very different standards. LLMs excel at generating the former, but they are nearly incapable of assessing the latter. Second, the abstraction problem, criticizing a method for not differing from "Transformers" as a whole, rather than from a specific architecture or training objective, makes rebuttals feel like arguing with a mirror. You cannot defend against a criticism that refuses to pin down its target. This is where the tool's weakness becomes the user's failure: the LLM does not know what it does not understand, and the human who pastes its output is effectively endorsing that ignorance.
This connects directly to a broader theme we have explored in our coverage of how LLMs navigate token spaces and expand technical fluency. In Exploring Paragraph Structure: How LLMs Navigate Token Space, we saw that these models process language as coordinate systems, not as grounded reasoning. That is precisely why they can propose a dozen "logically valid" criticisms without ever weighing their importance, they are operating in a space where plausibility is a substitute for evidence. Similarly, our piece on Expanding Your Tech Fluency emphasized that fluency with tools does not equal fluency with the underlying concepts. A reviewer who relies on an LLM to generate criticisms is exhibiting tool fluency, not domain mastery.
So what is the practical takeaway? If you are using an LLM to assist with reviews, treat its output as a *starting point for your own judgment*, not a draft to submit. Ask yourself: Would I raise this concern if I had not seen it generated? Does it attach to a specific method, a concrete number, or a clear causal mechanism? If the answer is no, delete it. And if you are on the receiving end of such reviews, do not feel obligated to address every speculative aside. Push back by asking for the specific prior work or the exact condition that would change the paper's conclusion. The real risk is not that LLMs produce bad reviews, it is that we stop noticing the difference between a review that is *comprehensive* and one that is *correct*. One open question worth watching: as these tools become more integrated into academic workflows, will journals and conferences start requiring reviewers to disclose their use, or will we simply accept a new normal where the burden of filtering always falls on the author? That decision will shape the integrity of peer review for years to come.