1 min readfrom Machine Learning

For the people who got reviews back from neurips, cvpr, eccv, etc and also tested their paper through an agentic reviewer like the stanford one, how different were the reviews? [D]

Our take

For those who recently received reviews from NeurIPS, CVPR, ECCV, or similar conferences, and also utilized agentic reviewer tools like the Stanford model, a compelling question arises: how do the reviews compare? We're exploring the divergence between human and LLM assessments, seeking insights into this evolving landscape. Early indications suggest significant variations, prompting a deeper understanding of how AI-assisted review impacts the peer review process. For further context on related challenges, see our article, "My Model Was Cheating on Its Own Test."

The recent Reddit thread questioning the divergence between human and LLM reviewer feedback on papers submitted to top-tier conferences like NeurIPS, CVPR, and ECCV highlights a fascinating and increasingly relevant challenge in the AI research landscape. The emergence of agentic reviewer systems, exemplified by the Stanford project, offers a tantalizing glimpse into automated peer review, but also raises critical questions about the quality and nature of the feedback received. It’s not simply about whether an LLM can mimic a human reviewer—it’s about *how* their perspectives differ, and what that means for the future of scientific evaluation. We’ve seen similar anxieties about data integrity and unexpected biases in model behavior, as explored in “My Model Was Cheating on Its Own Test,” which underscores the importance of rigorous evaluation even within seemingly automated processes. The ability of a model to subtly exploit weaknesses in a dataset is a reminder that automated systems aren't inherently objective.

The core of the question—what are the differences?—likely revolves around several factors. Human reviewers bring years of experience, intuition, and a nuanced understanding of the field's historical context and ongoing debates. They can assess the originality of an idea, its potential impact beyond immediate metrics, and its alignment with broader research trends in ways that current LLMs, while impressive, often struggle to replicate. LLMs, on the other hand, excel at identifying technical errors, inconsistencies in methodology, and areas where the writing could be clearer. They can systematically scan a paper for these issues with remarkable speed and accuracy. Furthermore, as discussed in “RAG Workflow and Loop Engineering: The Dispatcher That Decides When to Loop and When to Stop,” the reliability of any AI-driven assessment is fundamentally tied to the quality of the data and architecture used to train it – a point particularly relevant when considering the biases inherent in the datasets used to train large language models. The Reddit thread's inquiry also echoes concerns raised in “How much does adding an honest limitations section hurt the paper?” – a willingness to acknowledge weaknesses can be perceived differently by human and machine reviewers, potentially influencing overall scores.

The implications of these differences are significant. If LLM reviewers consistently identify technical flaws that human reviewers miss, it could lead to a more rigorous and error-free evaluation process. However, if they fail to grasp the broader significance of a contribution or provide feedback that is overly focused on superficial details, it could stifle innovation and discourage researchers from pursuing high-risk, high-reward ideas. The challenge lies in harnessing the strengths of both approaches—leveraging LLMs for their efficiency and technical acuity while retaining the judgment and contextual understanding of human experts. A hybrid model, where LLMs pre-screen papers and flag potential issues for human reviewers to address, seems like a promising avenue for exploration. Such a system could drastically reduce the workload on human reviewers, allowing them to focus on the most critical aspects of evaluation.

Ultimately, the feedback from this Reddit thread signals a necessary shift in how we think about peer review in the age of AI. It’s not about replacing human reviewers entirely, but about augmenting their capabilities and creating a more efficient and equitable evaluation system. As these agentic reviewer systems continue to evolve, it will be crucial to carefully monitor their performance, identify their biases, and develop strategies to mitigate their shortcomings. The question moving forward isn't simply whether LLMs can review papers, but *how* we can best integrate them into the peer review process to ensure that the most impactful and innovative research gets the recognition it deserves.

Hello,

I was curious about the differences you can get from the human reviewers and the llms.

Any insight is welcome, thank you!

submitted by /u/obliviousphoenix2003
[link] [comments]

Read on the original site

Open the publisher's page for the full experience

View original article