Teach neural networks when to say "I don't know" with HALO-Loss.

In the realm of neural networks, the challenge of distinguishing between uncertainty and false confidence is crucial.

3 min readMachine Learning

Neural networks are bad at admitting ignorance, and that is a problem we should have taken more seriously a long time ago. The HALO-Loss approach directly addresses this by giving models a built-in, mathematically sound way to say "I don't know." It is not a flashy trick or a band-aid; it is a structural fix to a fundamental flaw in how we train classifiers. The fact that it does this without sacrificing accuracy is what makes it worth your attention.

The practical implications are straightforward. Standard Cross-Entropy loss pushes features away from the origin indefinitely, which means the model has no natural place to represent uncertainty. HALO flips that by bounding confidence to a finite distance and attaching an "Abstain Class" right at the origin. For anyone working on safety-critical systems, this is the difference between a model that quietly fails and one that raises its hand. The numbers here are not marginal either: calibration error drops from roughly 8% to 1.5%, and false positives on far out-of-distribution data are cut by more than half. That is not a trade-off; that is an upgrade on both fronts.

What stands out is that this is not buried in a research paper behind paywalls. The author has open-sourced the code and written a breakdown of the math, including the practical pitfalls like dealing with high-dimensional Gaussian "soap bubbles." That matters because it lowers the barrier to entry. You do not need to be a machine learning researcher to test this on your own data. You can run it, see if your network's overconfidence drops, and decide for yourself. That is how progress should feel: accessible, measurable, and direct.

The real takeaway is that this is a viable path toward safer models without the usual tax. If you are training a CLIP-style model or any classifier where a wrong answer is expensive, HALO gives you a rejection threshold that is not bolted on after the fact. It is native to the training process. That is a meaningful step forward, and it is available now. Try it, stress-test it, and see if your model finally learns when to stay quiet.

From Machine Learning

Current neural networks have a fundamental geometry problem: If you feed them garbage data, they won't admit that they have no clue. They will confidently hallucinate. This happens because the standard Cross-Entropy loss requires models to push their features "infinitely" far away from the origin to reach a loss of 0.0 which leaves the model with a jagged latent space. It literally leaves the model with no mathematically sound place to throw its trash.

I've been working on a "fix" for this, and as a result I just open-sourced the HALO-Loss.

Read the original at Machine Learning