A dataset paper is accepted at a major computer vision conference. The authors promise the dataset in the paper, point to a GitHub link, and then deliver nothing. The repository is empty. It was always empty. The authors are not responding to emails. The conference requires the dataset to be available, but nobody at the conference checked, and now the person who noticed this has no clear path to file a complaint. That is the situation posted on Reddit, and it is not a small procedural annoyance. It is a breakdown of the trust model that makes peer review meaningful.
We spend a lot of time talking about model quality, but the foundation of any credible system is the data it is built on. If you have trained a model on a dataset that does not exist, you have not contributed science; you have contributed a claim. And when the community cannot verify that claim, the entire ecosystem weakens. This is the same problem we saw with Clean Data Starts With Catching AI Slop Before It Skews Your Model: garbage in, whether from sloppy scraping or broken promises, quietly poisons everything downstream. The difference here is that the slop is not a mistake; it is a structural failure in how we reward publication.
The person who posted this is not being unreasonable. They tried the authors first. They read the paper, clicked the link, watched the repository stay empty. They are not asking for special treatment. They are asking for the system to have a consequence for a clear violation of its own stated rules. And the fact that they have to ask who to contact is the real indictment. When a conference accepts a dataset paper, it is not doing the authors a favor. It is giving them a platform, credibility, and a line on their CV. In exchange, the conference is supposed to guarantee that the contribution is real. That guarantee is not being enforced.
This also connects to a broader issue we have seen in the field: the gap between what is published and what actually works in practice. In Exploring Real-World Computer Vision: Deployments, Edge Models, and Current Challenges, we discussed how models that look great in a paper often struggle when they hit real-world constraints. But at least in that case, the authors are usually trying to solve a hard problem. Here, the problem is not technical difficulty. It is accountability. And while tools like the Forrester Function can help us think more carefully about optimization, no amount of mathematical sophistication fixes a process that rewards claiming without delivering.
So what is our honest take? The person who posted this should not have to solve this alone. Conferences need a standing integrity desk, a clear email address, and a public process for handling these complaints. But until that exists, the burden falls on reviewers and area chairs to check the links in the papers they are approving. That is a small, concrete step that would have caught this immediately. If you are a reviewer, open the repository. If it is empty, reject the paper. That is not a radical idea. That is the least we can do to keep the literature honest. The next time you read a paper that promises a dataset, ask yourself if you would bet your next experiment on it. Because right now, the odds are not as good as they should be.