The machine learning community has a quiet culture of amnesia, and the editorial captures this perfectly. A generation of practitioners was raised on a set of rules that sounded like mathematics but were really just folklore wearing a lab coat. The bias-variance "bull's eye" diagram, the absolute prohibition on touching the test set, the belief that big models are inherently doomed, and the assumption that ensemble methods are always superior. Those were the guardrails. Then the field moved on, the guardrails proved to be cosmetic, and nobody issued a retraction. The authors of those theories just stopped teaching them and started riding the deep learning wave. This is not an attack on theory itself. It is an attack on the way theory was weaponized as a substitute for critical thinking. The uncomfortable truth is that many of those rules were derived from contrived examples that had nothing to do with real-world data, and they persisted because they were easier to test than to challenge. For anyone building systems today, the practical takeaway is that the field has swung hard into empiricism, but that does not mean we should abandon structure altogether. The question is not whether theory is dead, but whether we can tell the difference between a theorem and a habit. The editorial asks if anyone chooses an optimizer or a model because of a theoretical guarantee anymore, and the honest answer is mostly no. We use ADAM because it works on our loss surface, we use transformers because they scale, and we use ensembles because they squeeze out a few more points on a leaderboard. That is not a failure of rigor. It is a shift toward a different kind of rigor, one that is validated by reproducible results rather than elegant proofs. If you are a practitioner feeling guilty about not being able to justify every architectural choice with a proof, stop. The Clean Data Starts With Catching AI Slop Before It Skews Your Model piece shows how even simple filtering decisions can have a measurable impact on model accuracy, and nobody needs a theorem to justify cleaning your data. The deeper issue is that the field has not yet produced a new set of guiding principles to replace the old ones. We are left with a patchwork of heuristics, many of which are just as dogmatic as the theories they replaced. For example, the idea that you should never look at the test set has morphed into the practice of over-optimizing on a public leaderboard, which is arguably worse because it is a form of implicit overfitting that is harder to detect. And the belief that ensembles are always better ignores the cost of complexity, latency, and interpretability. A more mature stance is to treat every rule as a hypothesis to be tested in your specific context. That is what Exploring Real-World Computer Vision: Deployments, Edge Models, and Current Challenges demonstrates with mobile models, where theoretical assumptions about capacity often break down under real-world constraints. Similarly, Explore the Forrester Function: Beyond Mathematics, a Tool for Machine Learning reminds us that mathematical functions can still serve as useful benchmarks, but only if we remember they are tools, not truths. What we would tell a reader who asks whether theory still matters is this: theory is most valuable when it is generative, not prescriptive. It should suggest what to try, not dictate what to do. The moment a rule stops being questioned is the moment it becomes a liability. The real challenge is not finding a theory that explains everything, but building the judgment to know when a rule applies. That judgment is earned by breaking things, measuring the damage, and documenting what actually happened. The field deserves a new kind of folklore, one that is transparent about its failures and open to revision. The next time someone tells you that a model will not work because of a theoretical limitation, ask them to show you the data. And if they cannot, feel free to move on.
big data performance
Theory guided machine learning: A practice worth rediscovering.
The gap between machine learning theory and practice has never been wider, and the confusion is understandable.
4 min readMachine Learning
There was a period in the development of machine learning where application seemed to be informed by theory. Some of the best known theories include:
Most of these theories started out as mathematical statements (albeit on some contrived examples that have nothing to do with reality). At some point, these theories became folklores and were widely reproduced in textbooks and taught in classrooms, even making their ways into standard interview questions at data science related companies. Every student had to remember that bias-variance "bull's eye" diagram as if it was relevant in practice.