TMLR reached out to the authors of 10 papers slated for desk rejection, in an attempt to understand if the authors could explain the paper they submitted [D]
Our take
The recent experiment conducted by TMLR, reaching out to authors of papers flagged for desk rejection to gauge their understanding of their own work, is deeply concerning, and warrants serious reflection within the AI research community. The findings – a significant number of authors unable to answer basic or even high-level questions about their submissions – highlight a potential systemic issue within the rapid-publication environment of machine learning. It’s easy to dismiss these results as anecdotal, but the sheer volume of problematic responses – three unable to answer basic questions, three struggling with technical details, and one paper even containing a major flaw identified during questioning – suggests a pattern that deserves attention. This isn’t about individual failings; it's about the pressures and incentives shaping how research is conducted and presented. We’ve seen similar concerns surface recently, such as the issues surrounding hallucinated references highlighted in [NeurIPS Reference Check Response[D]], demonstrating the challenges in ensuring the integrity of published work. The speed at which models and techniques are being developed often outpaces the rigor of verification and deep understanding, and this TMLR experiment provides a stark illustration of that imbalance. The need for robust quality control is further emphasized by the ongoing conversations around potential pitfalls like those discussed in [Silent Broadcasting Can Ruin Your Model], which demonstrate how subtle errors in implementation can lead to significant performance issues that are difficult to detect.
The root causes are likely multifaceted. The sheer volume of submissions to top-tier conferences, driven by the competitive landscape and the desire for rapid career advancement, creates immense pressure on researchers. This can lead to rushed writing, superficial understanding, and a focus on novelty over thoroughness. The rise of large language models and AI-assisted writing tools, while offering benefits, also introduces the risk of authors relying on these tools to generate text without fully grasping the underlying concepts. Furthermore, the increasing complexity of machine learning models and techniques requires a level of specialized expertise that may be difficult to maintain across all areas of research. It’s not simply about having a brilliant idea; it’s about possessing the deep technical knowledge to rigorously validate and explain that idea, something that the TMLR experiment clearly showed was lacking in a significant number of cases. This contrasts with the innovative work being done in areas like voice simulation, as demonstrated by companies like Treble, who are focused on building and refining these technologies with a clear understanding of the underlying science [Iceland-based Treble raises $18 million for its voice simulation platform].
The implications of this are far-reaching. The credibility of AI research, already facing scrutiny, is further eroded when published work is found to lack a fundamental level of understanding. This impacts not only the academic community but also the broader ecosystem of AI development, where these papers often serve as the foundation for new products and services. A lack of rigor can lead to flawed implementations, unreliable models, and ultimately, a loss of trust in AI technology. The current peer-review process, while valuable, appears to be insufficient in catching these deeper issues. Simply identifying related work and demonstrating incremental improvements isn't enough; reviewers need to probe the fundamental understanding of the authors and their ability to defend their claims. The incentives within the system need to shift towards rewarding depth of understanding and rigorous validation, rather than simply rewarding novelty and publication volume.
Looking ahead, it's crucial to explore ways to strengthen the quality control mechanisms within AI research. This might involve more in-depth reviewer training, the introduction of oral examinations alongside paper submissions, or even the development of AI-powered tools to automatically assess the technical soundness of papers. More importantly, the culture within the field needs to change – a shift away from the relentless pursuit of publication and towards a greater emphasis on intellectual rigor and deep understanding. The question is not just *what* are we publishing, but *do we truly understand* what we are publishing, and can we adequately explain it? The TMLR experiment serves as a vital wake-up call, urging us to prioritize substance over speed and to ensure that the foundations of AI research are built on a bedrock of genuine understanding.
You can read more here: https://medium.com/@TmlrOrg/asking-authors-about-their-own-papers-3d2e04e5dee0
The results are (imho) concerning. Taken from the article:
Of the ten submissions:
- Authors of one paper withdrew their submission.
- Authors of one paper said they were unavailable due to other commitments.
- Authors of one paper scheduled a meeting but did not show up.
- Authors of three papers were unable to answer basic questions about the paper.
- Authors of three papers answered questions about high-level ideas in the paper but had difficulty when asked further questions on technical details.
- Authors of one paper answered all of my questions (although [the interviewer, the Co-EiC] identified a major flaw in that paper)
[link] [comments]
Read on the original site
Open the publisher's page for the full experience