radar object classification

Tracking radar objects over time unlocks deeper classification insights

A single radar scan of an object averages just 2.9 points, barely a whisper of data. Bruno Pinto's work shows that accumulating observations over a tracked object's history unlocks far deeper classification insights,…

3 min readMachine Learning
Tracking radar objects over time unlocks deeper classification insights
Multi scan radar object classification on RadarScenes [P]

The most instructive result in Bruno Pinto's radar classification experiment isn't the final accuracy number, it's the fact that simply pooling 20 scans of raw observations, with no attempt to model time at all, recovered 86% of the total improvement over a single-scan baseline. That finding matters far beyond autonomous driving. It suggests that for many real-world sensor problems, the bottleneck isn't sophisticated temporal architectures; it's giving the model enough data in the first place. And that has direct implications for how teams building autonomous systems should allocate their engineering resources, especially as companies like Waymo's Texas Fleet Grows Significantly, Reflecting Rapid Expansion scale their operations and need classification systems that work reliably at production volumes.

Pinto's work on the RadarScenes dataset is a clean ablation. A single radar scan of an object contains an average of just 2.9 points, barely enough to distinguish a pedestrian from a large vehicle. By accumulating observations over a tracked object's history, even without respecting scan order, the macro F1 score jumped from 0.7370 to 0.8613. Adding a causal GRU that does model temporal order pushed performance further to 0.8895, but the incremental gain was a relatively modest 0.0282. The takeaway is precise: temporal dynamics matter, but they matter less than simply having more observations. This is not a knock against sequence models; it is a reminder that representation quality at the per-scan level is the binding constraint. When Pinto tested larger GRUs, transformers, and state space models on the same frozen embeddings, performance landed in a narrow 0.86-0.89 band. The architecture choice for temporal aggregation was far less consequential than the decision to accumulate scans in the first place.

This has practical consequences for anyone building perception stacks. If your sensor data is sparse, and radar often is, the most cost-effective path to better classification may be to extend your tracking buffers rather than chase the latest sequence model architecture. That is a liberating finding for teams that lack the resources to fine-tune large transformer backbones. It also raises a question that Pinto's experiment doesn't answer: how many scans is enough? The 20-scan window here was a fixed choice; the plateau across architectures hints that a smaller window might have captured most of the gain, but we don't know where the diminishing returns curve bends. That is worth watching as autonomous vehicle fleets grow and the question of safe deployment becomes more urgent, a topic TechCrunch Mobility: How do we know when an AV is safe enough? explores in depth. For now, the concrete takeaway is this: before investing in complex temporal models, measure how much you can gain by simply looking longer. The answer may surprise you, and it will almost certainly save you time.

From Machine Learning

I built a radar object classifier on RadarScenes, extending a prior single-scan classifier to accumulate observations over a tracked object's history instead of classifying each scan in isolation.

A single RadarScenes object instance contains only about 2.9 radar points on average, very sparse. A single scan also can't capture temporal characteristics: RCS and micro-Doppler both vary continuously as an object moves. Pedestrians produce characteristic micro-Doppler from limb motion; different object classes show different RCS fluctuation patterns as aspect angle and scattering geometry change scan to scan. Accumulating observations gives both higher point density and provides temporal dynamics.

Read the original at Machine Learning