NeurIPS AI Assisted Review authors/reviewers? [D]
Our take
The recent Reddit thread discussing experiences with AI-assisted peer review at NeurIPS highlights a critical juncture in the evolution of academic publishing. The author’s observations – the disparity in review quality, the breach of double-blindness protocols, and the frustration with reviewers unfamiliar with established concepts – resonate with growing concerns about the integration of LLMs into traditionally human-driven processes. It’s a complex situation, amplified by the speed at which AI tools are being adopted. This isn't merely about the technology itself, but about the preparedness of the academic community to adapt their workflows and expectations. The situation mirrors broader anxieties explored in “A Mechanistic Explanation of Prompt Injection (and why you should study roles)”[/post/a-mechanistic-explanation-of-prompt-injection-and-why-you-sh-cmsnjjs3p08s3mi9z8el9ga39], underscoring the potential for unintended consequences when relying on AI without a robust understanding of its limitations.
The core issue, as the author points out, seems to be a lack of critical engagement with the tools themselves. The suggestion to explicitly acknowledge the potential for LLM assistance – “look, the point of an LLM assisted review is that if you don’t even know this material, you can ask it questions” – is a valuable one. It shifts the focus from simply using AI to generate text to leveraging it as a supplementary resource for deeper understanding. However, it requires a cultural shift within peer review. Currently, the implicit assumption is that reviewers possess a comprehensive knowledge base, and the use of external tools is either discouraged or considered a sign of inadequacy. Encouraging reviewers to transparently utilize AI as a tool for clarification and verification, rather than a replacement for their expertise, is essential. The uneven quality of reviews, as described, suggests that some reviewers are treating LLMs as shortcuts, while others are failing to recognize their potential for improving the review process. Relatedly, the exploration of causality within AI research, as discussed in "73 NeurIPS workshops, and not a single one on Causality [R]"[/post/73-neurips-workshops-and-not-a-single-one-on-causality-r-cmsnjk3fa08sbmi9zhmc62t87], highlights the need for a more rigorous, foundational understanding of AI's capabilities and limitations across disciplines.
The breach of double-blindness is particularly troubling. It demonstrates a lack of adherence to fundamental ethical principles and underscores the need for stronger safeguards. While the reviewer’s justification – identifying LLM-generated content – is understandable given the context, it sets a dangerous precedent. It suggests that the potential for AI detection outweighs the importance of maintaining anonymity, potentially introducing bias and undermining the integrity of the review process. Addressing this requires not only technical solutions (e.g., better AI-detection tools) but also clear guidelines and enforcement mechanisms to ensure compliance with ethical standards. Furthermore, the observation that reviewers sometimes focus on minor details rather than substantive issues, even in the absence of LLM assistance, highlights a pre-existing problem within the peer review system. The introduction of AI doesn't necessarily *create* these issues, but it can exacerbate them if not managed thoughtfully. The challenges faced in training models, as noted in "3 Collapsing models [R]"[/post/3-collapsing-models-r-cmsnjiiwe08qnmi9zr8xmpw6y], serve as a reminder that AI isn't a panacea and requires careful calibration and oversight.
Ultimately, the NeurIPS experience serves as a valuable case study in the challenges and opportunities of integrating AI into academic publishing. It’s clear that simply adopting LLMs without addressing the underlying cultural and ethical considerations is a recipe for unintended consequences. Moving forward, the focus should be on developing frameworks that empower reviewers to utilize AI responsibly, promote transparency in the review process, and safeguard the integrity of scientific inquiry. The key question now is: how can we design systems that leverage the potential of AI to enhance, rather than undermine, the crucial function of peer review in advancing knowledge?
Out of curiosity, if you were a reviewer or author, how did the review period go?
For me, it was weird, because I gave reviews with specific details (what specifically could have been better, how to fix it), but realized other reviewers gave similar superficial reviews. Even the paper which was a control for me (no LLM), I gave specific comments, but other reviewers focused on minor things.
During the discussion period for one paper, one reviewer broke the double blindness condition, and gave specific examples of what the LLM gave and justified their reject…..but they didn’t even state that in their initial review (nor engaged with the author rebuttals). There was no also no sense of: “author said this was unclear, check with the LLM to see what’s the issue”
For one of my own papers, we had great scores for originality and significance, but had low scores for clarity, with at least two reviewers finding difficulty understanding established notation and concepts, and I’m wondering whether it would have been better to break the double blindness and said: look, the point of an LLM assisted review is that if you don’t even know this material, you can ask it questions, like if other papers use the same notation, how our paper compares with them, etc…
[link] [comments]
Read on the original site
Open the publisher's page for the full experience