TMLR

When authors can't explain their own research, the process falters.

Ten authors.

3 min readMachine Learning

Ten authors submitted papers to a venue called TMLR. Ten. And when the editors asked those authors to explain their own work, only one person could answer every question, and that one paper still had a major flaw. Let that sink in for a moment. This is not a story about a single bad actor or a technical hiccup. It is a window into a systemic problem, and it should worry anyone who lives and works in the AI research community, whether you are a seasoned reviewer, a junior researcher, or a practitioner just trying to figure out which papers are worth your time.

The results are concerning, but they are not surprising to anyone who has been watching the field struggle with a flood of low-quality submissions. The editors of TMLR did something simple and radical: they asked authors to walk them through their own papers. One author withdrew. One was too busy. One scheduled a meeting and never showed up. Three could not answer basic questions. Three could answer high-level questions but fell apart on technical details. Only one author demonstrated genuine command of the material, and even that paper had a critical flaw. This is not a story about the limits of peer review. It is a story about a culture that rewards volume over understanding, and it has real consequences for how we build and trust AI systems.

For our readers, this should change how you approach the literature. When you read a paper, ask yourself a simple question: could the authors pass a basic interview about this work? If they cannot, what does that say about the reproducibility of the results? What does it say about the validity of the conclusions you are about to build a product on? This is not an abstract concern. As we explored in Exploring Real-World Computer Vision: Deployments, Edge Models, and Current Challenges, the gap between a polished paper and a working system is often vast. And when authors cannot explain their own methods, the gap becomes a chasm. The same lesson applies to the tools we use to analyze these systems, as discussed in Explore the Forrester Function: Beyond Mathematics, a Tool for Machine Learning, where the underlying assumptions matter more than the surface-level math.

This also raises uncomfortable questions about the incentives that got us here. We have seen repeatedly, as in ICLR Submissions Exposed: Addressing Data Privacy Concerns in AI Research, that the pressure to publish can lead to shortcuts that undermine the integrity of the entire enterprise. The TMLR experiment is a direct challenge to that culture. It suggests that we should not just ask whether a paper passes peer review. We should ask whether the authors can stand behind their work in a live conversation. That is a higher bar, but it is also the bar that separates real research from noise. The concrete point to watch is this: if venues like TMLR continue this practice, and if they start publishing the results, authors will be forced to adapt. Some will withdraw, as one did here. Others will actually learn their papers inside out. And the ones who cannot will not be able to hide behind the PDF anymore. That is a future worth pushing for, and it starts with demanding more from the people who sign their names to the work.

From Machine Learning

You can read more here: https://medium.com/@TmlrOrg/asking-authors-about-their-own-papers-3d2e04e5dee0

The results are (imho) concerning. Taken from the article:

Read the original at Machine Learning