confidence scoring

Small models, smarter labels: confidence scoring without the hype

Small classification models that score confidence over a list of labels, rather than generating free text, deserve more attention.

3 min readMachine Learning

There is a quiet revolution happening in applied machine learning, and it looks nothing like a demo video. It looks like someone spending the time to rewrite label descriptions and watching accuracy jump twelve points. Gokul Krishna ran his own evaluations of two small classification models, TypeSafe AI's hosted Jev and the open-source Laya, in one of the most grounded pieces of practical AI writing we have seen in months. It deserves attention not for its headline numbers, but for what it reveals about how this technology actually works when you stop chasing benchmarks and start solving real problems.

The headline finding is that describing labels by what the input looks like, rather than by the intent behind it, improved accuracy from 65% to 77% using the same model. That is a 12-point gain from a text edit, not a model update. It cuts directly against the hype cycle that tells users the answer is always a bigger, better, more expensive model. We have covered Jev's typed probabilities before in Jev delivers typed probabilities, not text, for faster data decisions, and we have seen the Intern-Decision family outperform it at certain sizes. But a different point emerges: the interface between the model and the human matters as much as the architecture. If you describe an AML transaction by what the criminal is trying to achieve, you are asking the model to read minds. Describe it by what the transactions look like, and you are asking it to read data. That shift is not about hype; it is about discipline.

The calibration data is equally telling. Jev's expected calibration error of 0.013 allowed a 90% confidence threshold to raise accuracy from 92.3% to 97.6% while still covering 82% of cases. The untuned Laya, with an ECE of 0.486, dropped 7% of questions for only a 2-point gain. We previously reported that Jev's confidence accuracy jumped 68% after alignment, and these results reinforce that calibration is not a nice-to-have; it is the mechanism that makes thresholds operational. A model that cannot tell you when it is guessing is not a decision model; it is a roulette wheel with better marketing.

The most honest finding is that some tasks have no signal. Classifying a single transaction as laundering or not gave 54% accuracy, essentially chance. That is a limit of the task, not of the model. Too many vendors and practitioners treat AI as a universal solvent. It is not. Some problems simply do not contain enough information to solve at the granularity you want. Knowing when to fall back, when to say "this needs a human" or "this needs more data", is itself a capability that separates mature systems from prototypes.

What we want to see next is whether the per-task threshold tuning Krishna describes can be automated without sacrificing the calibration gains. That is the open question that will determine whether these small, confidence-scoring models become practical tools for enterprise workflows or remain research curiosities. The answer will not come from a press release. It will come from more evaluations like this one, run by people who care more about what works than what sounds impressive.

From Machine Learning

I ran my own evals of two small classification models that score confidence over a list of candidate labels instead of generating text: TypeSafe AI's hosted Jev, and Laya, an independent open-source alternative.

Some findings that I think apply beyond these two models:

Read the original at Machine Learning