AAAI 2027

Why Reproducibility Must Be More Than a Checklist Promise

Four papers for AAAI 2027, and not a single line of code or data to inspect.

4 min readMachine Learning

Four papers. Four empirical claims. Zero lines of code. That is the reality of one reviewer's AAAI 2027 batch, and honestly, it should give anyone in this field pause. The original poster is not asking whether missing code is ideal; they are asking whether it should be a hard reject when the rules clearly say it should be provided. We think the more useful question is why we keep having this conversation at all. The checklist exists, the policy is stated, and yet here we are, weighing whether "trust me, it works" is a sufficient substitute for verifiable artifacts. This is not about policing authors or assuming bad faith. It is about the quiet erosion of accountability that happens when we let the absence of evidence become a negotiable point. For a community that prides itself on rigor, we are remarkably comfortable with the phrase "we will release it after acceptance," as if the review process is a formality and the real work happens later. It is not, and it does not.

The practical tension is real, though. Reviewers are stretched thin, and auditing code is time-consuming. The thread the poster references even makes that point: someone who helped write the checklist argued that reviewers rarely have time to audit code anyway. There is truth there, but it is a truth that indicts the system, not the standard. If we accept that we cannot verify, we should at least be honest about what we are doing. We are reviewing narratives, not results. We are evaluating the story the paper tells, not the data behind it. That is a meaningful distinction, and it should change how much confidence we assign to a paper's conclusions. The poster is right to flag it. They are right to ask for anonymized code in the rebuttal. But here is where we would push back: do not wait for the rebuttal to make the point. If a paper's entire weight rests on empirical numbers, and those numbers come with no way to check them, that is not a minor oversight. It is a fundamental gap in the contribution. It should lower your confidence score, and it should lower it significantly.

This connects to a broader pattern we have been watching across the AI research landscape. When we covered Unlock LLM Training: A Practical Guide to Distributed Algorithms, we noted how much of the field's progress depends on shared infrastructure and reproducible baselines. And when we looked at Exploring Paragraph Structure: How LLMs Navigate Token Space, the same lesson surfaced: understanding the mechanism matters less when you cannot see the data that produced the claim. The issue is not unique to AAAI, and it is not unique to this reviewer. It is systemic. We are rewarding papers that tell good stories with impressive numbers, and we are doing it while pretending that the absence of code is a minor logistical detail rather than a core flaw in the scientific method being applied. The Neurosurgery Match Requirements Highlight Growing Pressure on Medical Students piece we ran earlier showed a similar dynamic in a different field: when the bar for entry rises without a corresponding rise in accountability, the people who suffer are the ones who play by the rules.

So what is the takeaway here? It is not that missing code should be an automatic reject in every case. Funding constraints and IP concerns are real, and we should not pretend otherwise. But those exceptions should be rare, documented, and justified. The default should be that empirical claims come with the artifacts that support them. And when they do not, the review should say so plainly, not as a suggestion for future work, but as a condition of the paper's credibility. The poster asked how others are handling this round. We would say this: treat missing code the way you would treat a missing control group. It is not automatically disqualifying, but it demands a much higher burden of proof for everything else. If a paper cannot show its work, the numbers are just a story. And we have enough stories. What we need is evidence.

From Machine Learning

I got my batch of four papers for AAAI 2027. All four papers make empirical claims, none include code, data, or anything I can actually check. Just the PDF and the checklist. AAAI-27's own rules say code/data should be provided at submission, and "we'll release it after acceptance" doesn't count as reproducibility.

Read the original at Machine Learning