categorical labelling

Teaching Machines to Score from Categories, Not Numbers

When categorical labels are all you have, deriving a continuous score might feel like pulling a number from thin air.

3 min readTowards Data Science
Teaching Machines to Score from Categories, Not Numbers

The gap between the labels we have and the scores we actually need is one of the quiet frustrations of applied machine learning. Most teams have plenty of categorical data, but the continuous, fine-grained signals that drive real decisions often don't exist. Deriving a continuous score from categories by using low-capacity networks is a refreshingly practical look at a problem that usually gets buried under jargon about transfer learning or pseudo-labelling. It doesn't promise magic. It walks through the maths of why a deliberately constrained model can force the categories to encode a meaningful order, even when no one ever handed you a numeric target. That is a genuinely useful insight, because it shifts the burden from collecting more data to thinking harder about the structure you already have.

This matters more than ever when you consider how messy real-world data has become. We recently wrote about how Clean Data Starts With Catching AI Slop Before It Skews Your Model, and the point there applies directly here. If your categorical labels are noisy, or worse, contaminated by automated systems, then the continuous score you derive is only as trustworthy as the categories you started with. The approach is elegant precisely because it doesn't ask you to invent a score out of thin air. It asks you to trust that the relative ordering of categories carries information, and then uses a low-capacity network to extract that ordering without overfitting to the noise. That is not a trivial distinction. It is the difference between building a model that generalises and one that simply memorises the quirks of your training set.

There is also a human element here that the technical write-up only hints at, and it's worth spelling out. The same week we ran a piece on Talking to My AI Clone Taught Me to Question the Tech, which explored how quickly we anthropomorphise outputs and assume intent where there is only optimisation. The technique is a useful counterweight to that tendency. It is a reminder that sometimes the most responsible thing you can do with a model is not to make it bigger or more complex, but to deliberately limit its capacity so it can only learn the signal you actually care about. That is a discipline most teams skip. They reach for a larger model, more layers, more parameters, and then wonder why the scores feel off. Low-capacity networks are not a limitation here. They are the point.

What we would tell a reader is simple: do not wait until you have perfect labels to start building. Try the categorical approach first, but check your categories for contamination before you trust the output. The one specific takeaway worth quoting is this: a continuous score derived from categories is only as good as the constraints you place on the model that learns it. The moment you stop treating low capacity as a bug and start treating it as a feature, you open up a path to solving problems that previously felt impossible. Watch for the next time someone tells you they lack the data to train a scoring model. They might just be missing the categories they already have.

From Towards Data Science

A walkthrough of and the maths behind using low-capacity networks to acquire fine-grained scoring when only categorical labelling is available for training

The post Estimating from No Data: Deriving a Continuous Score from Categories appeared first on Towards Data Science.

Read the original at Towards Data Science