AAAI

Code submissions at AAAI 2027 raise questions about reproducibility standards.

A reviewer for AAAI 2027 is noticing a troubling pattern: a surprising number of papers arrive without any code.

4 min readMachine Learning

A reviewer for AAAI 2027 recently shared a frustration that should resonate across the entire research community: a surprising number of submissions arrive without any code implementation. This is despite the conference's explicit emphasis on reproducibility. The reviewer notes that they have always submitted their code, finding it leaves a strong impression, and they see no valid excuse for omitting it, especially when AI assistants can generate plausible empirical papers in a matter of hours. This is not a minor grievance; it is a signal that the foundation of our collective trust in published results is eroding.

We understand the counterarguments, of course. Concerns about idea theft, the effort of cleaning up code for public release, or the pressure to hit a deadline are real. But the reviewer's point cuts through those justifications: the barrier to entry for producing fake results has never been lower, and the absence of code makes it far easier to hide behind a convincing narrative. This isn't about punishing carelessness; it's about protecting the integrity of the scientific record. If we accept a paper without code, we are implicitly accepting that its results might be unverifiable. That is a dangerous standard to normalize. As we have seen in related discussions on Clean Data Starts With Catching AI Slop Before It Skews Your Model, the quality of our data and the tools we use to validate it directly impact the soundness of our conclusions. The same logic applies here: code is the data of empirical research.

Our take is straightforward: code submission should not be a point of praise; it should be a mandatory condition for review. If a paper makes empirical claims, the burden of proof rests on the authors to provide the means for those claims to be tested. We would tell the reviewer that their instinct to factor missing code into their initial scores is not just fair, it is necessary. They are not being punitive; they are being a diligent steward of the field. The practical consequence is that we all need to recalibrate our expectations. When you encounter a paper without code, the default assumption should be skepticism, not benefit of the doubt. The same way we would question a paper with missing data, we should question one with missing implementation. This is about more than one conference cycle; it is about setting a precedent. Consider the broader implications for how we build on each other's work, as highlighted in Exploring Real-World Computer Vision: Deployments, Edge Models, and Current Challenges, where practical deployment often reveals the gaps between a paper's claims and its actual utility. Code is the first step toward bridging that gap.

The reviewer's observation also points to a deeper cultural issue within AI research. We are in a race to publish, and the pressure to produce novel results can overshadow the need for rigor. The fact that an AI can write a fake empirical paper in hours is not an argument for ignoring the problem; it is an argument for building stronger verification mechanisms. We should not be asking if a paper *could* be fake, but rather assuming it might be until proven otherwise. The open question for our readers, then, is this: what are you doing to ensure your own submissions are models of transparency? The specific detail to watch is whether conferences like AAAI will move from encouraging code submission to enforcing it, because that is the only real deterrent. Until then, the onus is on reviewers to hold the line, and on authors to recognize that their code is not an optional extra. It is the very currency of trust.

From Machine Learning

I am now reviewing a bunch of papers for AAAI 2027 and it has surprised me the low amount of submissions with no code implementation. I don’t know if it has been only in my batch or it is common, but I was expecting very detailed appendices + code submission since AAAI is very explicit with the topic of reproducibility. I was planning to take this into consideration when assigning my initial scores, but I would like to hear your opinions. I have always submitted my code: it gives a very good impression and after reviewing process finishes we just…

Read the original at Machine Learning