reproducibility

Reproducible research needs more than promises: it needs code.

Twelve papers reviewed this year, and only one came with full code.

4 min readMachine Learning

Twelve papers reviewed this year. One with full code. Seven with none. That is not a slow slide into bad practice; it is a structural failure of the review process. The anonymous reviewer who posted this account from the NeurIPS season has given us a snapshot of a system that rewards hiding more than it rewards rigor. And the most damning detail is not the missing code itself. It is the incentive structure underneath: releasing code only increases the odds that a reviewer will find a bug and reject the paper. So the rational move, for anyone playing that game, is to hide the work. We have built a culture where obscurity is a strategy. That is not science. That is risk management.

The pattern should feel familiar to anyone who has watched Clean Data Starts With Catching AI Slop Before It Skews Your Model. There, the problem was noisy training data silently degrading a sentiment model. Here, the problem is silent code. Both are examples of the same underlying tension: the thing that makes a model trustworthy is also the thing that makes it vulnerable to scrutiny. When reviewers are the only line of defense, and when those reviewers are expected to find bugs in thousands of lines of code without running it, the system breaks. We saw it in the numbers: of the five papers that did share code, three contained obvious bugs that invalidated the results. Obvious bugs. Not obscure edge cases. That is not a quality gap. That is a signal that the authors were not even running their own code. The link to Exploring Real-World Computer Vision: Deployments, Edge Models, and Current Challenges is direct here. When you optimize for mobile deployment, you learn quickly that a model that cannot run is worthless. The same logic should apply to a paper. If the code cannot reproduce the results, the paper is a press release.

So what do we actually do about it? Desk rejection when code is missing is argued for. That is the right instinct, but it needs to go further. Desk reject the paper, yes. But also publish the reviewer reports that caught the bugs. Make the cost of hiding higher than the cost of sharing. And for the community, we should stop treating code as an appendix. It is not a supplement. It is the evidence. A paper without code is a claim without evidence. In any other domain, we would call that an opinion. The practical takeaway for anyone reading this: if you are a reviewer, demand the code. If you are an author, release it. And if you are a conference organizer, make the penalty for hiding it explicit and immediate. Otherwise, we are just curating a collection of abstracts that look good in a PDF and collapse on first contact with reality. The next step is to watch whether the major conferences adopt a code-required policy for the next cycle. That is the detail to track. Not because it will fix everything, but because it will tell us whether the field finally believes its own claims about reproducibility.

From Machine Learning

As review season for NeurIPS wraps up, I have now reviewed for 3 major conferences this year. And I'm noticing a worrying trend:

Out of the 12 papers I reviewed this year, only 1 provided full code (that runs the whole training pipeline from input dataset to output AUROC). 4 provided partial code with fragments of their method, but no ability to run the experiment end to end. And 7 provided no code.

Read the original at Machine Learning