performance regression detection

When Anomaly Detection Meets Sparse Data: Rethinking Performance Regressions

Evaluating anomaly detection with only ten healthy samples is a tightrope walk, and the questions here show a solid grasp of the risks.

3 min readMachine Learning

The tension in this setup is familiar to anyone who has tried to turn a handful of healthy runs into a reliable anomaly detector. You have ten samples per counter group, you are using leave-one-out to set the threshold, and you are holding regression runs aside as your test set. The instinct to protect the model from contamination is sound. The confusion about whether you also need a conventional train/validation/test split, or whether regression samples can stand in for unseen data, is exactly the right question to be asking before you finalise anything.

Here is the honest take: with ten healthy samples, a 60/20/20 split is not a rigorous evaluation strategy, it is a recipe for high variance. You would be training on six points, validating on two, and testing on two. That does not tell you anything stable about false-positive rates. Leave-one-out is the better call here because it uses all ten points for threshold selection while still giving you an estimate of how the model behaves on unseen healthy data. The regression samples can serve as a secondary check, but they are not a substitute for a proper healthy test set. They tell you whether the model catches the anomaly you already know about, which is useful, but they do not tell you how often the model cries wolf on normal variation. That distinction matters more than most people realise.

What you actually need is a second independent healthy dataset. Not a bigger training set, not more regression runs, but a fresh batch of normal operations that were never touched during development. This is where the evaluation shifts from model fitting to operational trust. You are not predicting a continuous value here, so MSE and MAE are the wrong lenses. You care about false-positive rate and detection rate. Those two numbers, reported together, tell you whether the threshold you set on ten samples will hold up when the system runs in the wild. The related conversation on Clean Data Starts With Catching AI Slop Before It Skews Your Model shows what happens when you filter based on a model you have not validated properly: you make confident decisions on weak evidence and pay for it later. The principle carries over directly. If you cannot trust your evaluation setup, you cannot trust the threshold, and if you cannot trust the threshold, you cannot trust the alerts.

The practical takeaway is this: keep leave-one-out for threshold selection, treat the regression runs as a smoke test rather than a final verdict, and invest the effort in collecting that second healthy dataset before you call the system done. It does not need to be large, ten more runs would already strengthen your confidence considerably, but it needs to be independent. The question worth watching is whether you can resist the temptation to keep tuning the threshold after you see the false-positive rate on that new data. If you tune on it, you have contaminated it, and you are back to square one. Build the separation into your workflow now, and you will have an evaluation that actually tells you something when it matters.

From Machine Learning

I’m working on performance regression detection using machine learning/anomaly detection.

Just trying to make sure the evaluation setup is correct before I finalise it.

Read the original at Machine Learning