The conversation around reproducibility in machine learning has reached a point where we need to stop pretending the problem is merely technical. When a researcher in physical AI has to trust a demo video because the lab setup costs more than their entire department's budget, we have crossed a threshold. Demos are curated performances. The incentives to show only the working segments are not a moral failing; they are a rational response to a system that rewards attention over verification. This is not about bad actors. It is about a structural shift where the cost of validation has become prohibitive for most, and the gap between what is claimed and what can be checked is widening by the quarter.
This connects directly to the challenges we have seen in Exploring Real-World Computer Vision: Deployments, Edge Models, and Current Challenges. In that discussion, the focus was on the messy reality of deploying models in uncontrolled environments, where the gap between a benchmark and a street corner is enormous. The same logic applies here. When a lab claims a model handles occlusion or lighting changes with a certain accuracy, the only honest way to verify is to run it yourself. But if you lack the proprietary hardware or the specific data pipeline, you are left with a number and a promise. The problems being solved are often so vague that the figures become meaningless. This is not a conspiracy. It is a natural consequence of a field that rewards novelty and marketing over the slow, unglamorous work of verification.
There is also a quieter, more uncomfortable issue at play: the reluctance to share code. Emailing authors and receiving silence is a common experience. We have all been there. The reasoning is rarely about protecting trade secrets. It is about protecting a competitive edge in a crowded field. If you have spent months tuning a model, handing over the exact configuration feels like giving away your lunch. This is not unique to machine learning, but the pace of the field amplifies it. The author contrasts this with the Manhattan Project or Apollo, where internal reproducibility was high even if outside verification was low. That comparison is useful, but it breaks down because those efforts had a single, measurable goal. Most ML research does not. It is exploratory, subjective, and often evaluated on metrics that do not capture real-world utility. You cannot mathematically check a claim that "the model understands context better" if the evaluation is a demo with cherry-picked examples.
So what do we tell a reader who asks whether reproducibility is dead? We tell them it is not dead, but it is becoming a luxury good. Only well-funded labs or large companies can afford to verify claims independently. The rest of us are left to triage: trust the source, look for reproducible baselines, and focus on problems where the evaluation is concrete. If you are working on a task with a clear metric, you have a fighting chance. If you are working on something vague, demand better evaluation before you invest. The specific thing to watch is the next wave of tools claiming to solve "general" problems. Until they provide open, verifiable baselines, treat the numbers as advertising. We would tell any reader this: do not abandon the principle, but do not be naive about the economics. The future belongs to those who can show their work, not just those who can show a video.