Correcting Probability Shifts When Under-Sampling Imbalanced Data

When building a predictive model for a heavily imbalanced dataset, particularly with a binary outcome, it's crucial to address how under-sampling impacts your logistic regression predictions.

3 min readData Science

The question of whether to rescale probability predictions after under-sampling is one of those deceptively simple issues that trips up even experienced modelers. The answer is yes, you absolutely need to correct for the sampling shift, but not because the raw probability scale matters in isolation. What matters is that your model's predicted probabilities are now systematically biased upward for the minority class, and that bias will distort any threshold you choose unless you adjust for it.

Here's the practical reality: when you under-sample the majority class to fit your data into memory, you're changing the prior distribution of your outcome. Your logistic regression will happily learn from this skewed sample, producing predictions that reflect the artificial 50/50 balance you created rather than the true 0.1% or 1% event rate in your full dataset. If you then pick a threshold based on those distorted probabilities, say, 0.5 as a naive default, you'll end up classifying far too many records as positive. The scale isn't arbitrary; it's just wrong. The fix is straightforward: you can apply a correction factor using the known sampling fractions, or you can simply recalibrate your probabilities on a holdout set that reflects the real class distribution. Either approach works, but skipping the correction means your threshold selection is built on a foundation that doesn't match the problem you're actually solving.

What this means for you, the person wrestling with a massive imbalanced dataset, is that your workflow needs to separate two distinct goals. First, you're under-sampling to make computation feasible, that's a legitimate engineering constraint, not a modeling choice. Second, you're building a classifier that will operate on real-world data where the base rate is what it is. These two steps don't have to fight each other, but they do require intentionality. If you only care about ranking (say, for an AUC-style evaluation), the raw predicted probabilities might be fine because rank order is preserved under monotonic transformations. But the moment you need to pick a decision threshold, to send a notification, approve a claim, or flag a transaction, you must correct the probabilities first. Otherwise, you're optimizing a threshold against a distorted version of reality, and your model's performance in production will quietly disappoint you.

So here's the concrete takeaway: build your model on the under-sampled data, but reserve a separate validation set that mirrors the true class distribution. Use that set to calibrate your predicted probabilities, whether through a simple intercept adjustment or a more flexible recalibration method, and only then choose your threshold. This isn't an optional refinement; it's the difference between a model that looks great in your training script and one that actually works when it meets your data. The scale isn't arbitrary, and pretending otherwise is how good models go bad.

From Data Science

I'm building a predictive model for a large dataset with a binary 0/1 outcome that is heavily imbalanced.

I'm under-sampling records from the majority outcome class (the 0s) in order to fit the data into my computer's memory prior to fitting a logistic regression model.

Read the original at Data Science