scikit-learn

How a single bug fix sharpens BayesianRidge uncertainty estimates

A subtle bug in BayesianRidge's uncertainty calculation slipped into sklearn 1.8, and this notebook walks through the exact formulas behind `predict` on both versions. It's a sharp exercise in tracing what changed, and…

4 min readMachine Learning

A bug in BayesianRidge uncertainty estimation got fixed in scikit-learn 1.9, and someone was kind enough to trace the actual `predict` code on both versions side by side. The notebook walks through exactly what changed, inviting readers to spot the difference before the reveal. This is the kind of work that rarely gets celebrated, but it matters enormously for anyone who relies on uncertainty numbers to make decisions. If you have ever used a confidence interval from a fitted model to decide whether to deploy it, to trust a prediction, or to flag an outlier, then this fix is not a footnote. It is the difference between a model that quietly lies to you and one that gives you an honest estimate of what it does not know.

The deeper lesson here is about the gap between what a library promises and what it actually computes under the hood. We have written before about how Clean Data Starts With Catching AI Slop Before It Skews Your Model, and this scikit-learn episode is the same story from the other direction. Garbage in, garbage out is a well-known risk for training data, but less attention goes to the silent corruption that happens inside a mature open-source library when a formula drifts from its intended derivation. The maintainers did not change the API. They did not add a feature. They just corrected a calculation, and anyone who upgraded without reading the release notes might never know that their previous uncertainty estimates were subtly wrong. That is the uncomfortable truth about software trust: it is not enough to verify the inputs and outputs. You have to verify the math in between.

What makes this notebook particularly useful is that it is not a rant or a conspiracy theory. It is a careful, reproducible comparison of two versions of the same function, showing the exact lines that changed and the practical impact on the resulting predictions. That is the right way to hold a library accountable. It is also a reminder that Exploring Real-World Computer Vision: Deployments, Edge Models, and Current Challenges and Explore the Forrester Function: Beyond Mathematics, a Tool for Machine Learning both touch on the same underlying theme: the gap between theoretical elegance and operational reality. A beautiful objective function or a state-of-the-art architecture means nothing if the implementation drifts from the derivation. The notebook turns that abstract concern into a concrete, inspectable artifact.

If a reader asked us what to do with this, we would say two things. First, do not assume your uncertainty estimates were ever correct just because the library was popular. Second, build a habit of regression-testing your models against known formulas, especially after a minor version upgrade. The scikit-learn team deserves credit for fixing the bug, but the real takeaway is that you are the last line of defense for your own analysis. A single notebook like this one can save you from a quiet, persistent error that would have corrupted every downstream decision you made based on faulty confidence. That is not a hypothetical risk. It is a bug that existed in a widely used library for years, and it was only caught because someone took the time to trace the code and share what they found.

From Machine Learning

sklearn 1.9 fixed a bug in how BayesianRidge computes its uncertainty. We traced predict on 1.8 and 1.9 and compared the two formulas it actually computes, see if you can spot what changed before the notebook tells you.

https://github.com/aadya940/scikit-verify/blob/master/examples/sklearn_bug_hunting.ipynb

Read the original at Machine Learning