BIRADS detection

Balancing the Unbalanced: Smarter Loss Functions for Medical AI

Three models collapsing toward BIRADS 1 is a familiar wall when a dataset leans that hard.

4 min readMachine Learning

A model that collapses into the majority class is not failing because of a mysterious bug. It is telling you something direct about the data and the objective you have chosen. When three separate attempts all land on the same outcome, the common denominator is usually the loss landscape, not the architecture. Cross-entropy with class weights is a reasonable starting point, but it assumes the weighting scheme actually rebalances the gradient signal. In a heavily skewed set like VinDr, where BIRADS 1 dominates, the model can satisfy the weighted loss by becoming confidently wrong on the minority classes while keeping the overall loss low. That is not a failure of effort; it is a signal that the loss function and the data distribution are not aligned with the task you actually care about.

This is a familiar pattern for anyone who has worked with medical imaging or any domain with extreme class imbalance. The instinct to add center loss is understandable, it tightens feature clusters, but it does not solve the root problem of representation. If the model never sees enough examples of BIRADS 2 through 5 to learn their boundaries, center loss will only make the collapse more precise. What you need is not a different loss function in isolation, but a strategy that changes how the model perceives the minority classes. This could mean oversampling those examples, using synthetic augmentation, or switching to a loss that explicitly optimizes for ranking or margin, like a triplet loss or a variant of focal loss. The choice depends on whether you want the model to be better at distinguishing rare cases or better at calibrating confidence. Both are valid, but they lead to different trade-offs.

We would also push back on the assumption that the wrong loss function is the primary suspect. The more likely issue is that the model is underfitting the minority classes because they are too few and too similar to the majority. In practice, we have seen teams spend weeks tuning losses when the real fix is to revisit the data split, check for label noise in the minority classes, or use a pretrained backbone that already understands the visual features of mammograms. If the model cannot separate BIRADS 1 from BIRADS 2 because the images are nearly identical, no loss function will save you. This connects to a broader lesson in model development: the loss function is only one lever, and its effect is bounded by the quality and structure of the data.

If you are stuck, we would suggest a simple experiment. Take a small balanced subset of the training data, say 200 examples per class, and see if the model can overfit it. If it cannot, the problem is capacity or optimization, not imbalance. If it can, then the imbalance is the bottleneck, and you should focus on sampling and augmentation rather than loss design. Also, consider evaluating on a metric that is not accuracy. For BIRADS, a weighted kappa or a one-vs-rest AUC will show you whether the model is actually ranking classes correctly, even if its hard predictions are biased. That is the metric a radiologist will care about. The takeaway is this: do not chase a new loss function yet. Diagnose the data first, and let the model's behavior on a balanced sample tell you where the real bottleneck is. That will save you more time than any single algorithmic tweak.

From Machine Learning

Trying to train 3 models for birads detection using cross entropy and center loss + class weights but all of them seem to collapse between birads 1 as the dataset (VinDr) im using is heavily unbalanced towards it, Would like to ask for input and opinion on what seems to be the case, am I using the wrong loss function?

Read the original at Machine Learning