rows.com

Exploring how an older model still competes on accuracy and speed today

CABiNet, a 2021 architecture that went quiet after ICRA, is back, and it's beating a 2026 generalist model on aerial segmentation.

4 min readMachine Learning
Exploring how an older model still competes on accuracy and speed today
CABiNet (ICRA 2021) vs YOLO26-sem on UAVid: accuracy, compute, and GPU latency [P]

The most honest thing about this benchmark is its author showing his work before making a single claim. The person behind CABiNet, a 2021 ICRA paper that went quiet, came back, rebuilt the repo, and asked a question most researchers avoid: how does a purpose-built efficient architecture from four years ago actually compare to a 2026 general multi-task model on the same aerial dataset? The answer on UAVid is not a blowout, but it is a clear statement. CABiNet-Large hits 67.14 mIoU at 4.44 milliseconds, while YOLO26x-sem, nearly 4.5 times the parameters and 8 times the FLOPs, lands at 64.41 mIoU and 13.09 milliseconds. The smaller CABiNet-S matches YOLO26s at nearly identical compute, 44.1 versus 44.4 GFLOPs, and beats it by 3.6 mIoU. That is not a universal win, and the benchmark's author says so plainly. On VDD and AeroScapes, YOLO26 pulls ahead, and the higher-end YOLO variants are faster at the low-accuracy end. But on UAVid, the dataset the original paper targeted, the older architecture holds the high-accuracy frontier. This matters because it challenges the reflexive assumption that newer means better across the board, especially when you are choosing a model for a specific deployment rather than a leaderboard chase.

What makes this comparison worth your attention is not the headline number but the discipline of the setup. The author standardizes the data representation, class weighting, and evaluation protocol, while letting each model keep its native training recipe. That is a defensible choice, and he invites the criticism directly. YOLO26 starts from Cityscapes and ADE20K pretraining, while CABiNet only has ImageNet, a clear asymmetry that likely favors the newer model, yet the older one still wins on the target metric. The per-class breakdown is even more telling. The gap comes almost entirely from small, thin classes: humans at +7.2 IoU, static cars at +5.9, moving cars at +5.1. On big region classes like building and road, the two are within half a point. This is not a story about one model being universally smarter. It is a story about where each model pays attention. CABiNet tracks the parking lot structure and the pedestrian; YOLO26 collapses them into a single mass. For anyone working on aerial imagery, where small objects are the whole point, that is not a trivial detail. It is the difference between a model that sees a scene and one that sees a smudge.

The practical takeaway for our readers is sharper than "old beats new." It is that FLOPs are a poor proxy for real-world latency on GPU, and that accuracy per millisecond at the higher-accuracy end is where the real decisions get made. MobileNetV3's depthwise convolutions are cheap on paper but not on Tensor Cores, which is why CABiNet's latency is not as low as its FLOPs suggest, and why YOLO26n and s legitimately own the low-latency corner. If you need 400 FPS, do not look at CABiNet. If you need the best accuracy you can get under 5 milliseconds on a single consumer GPU, this benchmark says a 2021 architecture with a clean training recipe is still your answer. That is a specific, actionable claim you can take to your next model selection meeting. The author also flags the limits: single run per config, no seed sweep, no statistical significance on the sub-point differences, and no end-to-end deployment cost for 4K source frames. Those are real caveats, not footnotes.

What we would tell a reader who asks us whether this means you should switch your aerial segmentation stack tomorrow is this: do not switch on mIoU alone, but do not dismiss the result because the paper is old. The open question the author poses is the one worth carrying forward. If you standardize data and evaluation but leave each model's training recipe intact, are you comparing architectures or recipes? He does not have the answer, and neither do we. But he has given us a reproducible starting point, and that is more than most benchmark posts ever offer. Watch whether anyone runs the seed sweep or extends this to a second aerial dataset with a matched pretraining scheme. That is the experiment that will actually settle the argument.

From Machine Learning

Disclosure up front: I'm the original first author of CABiNet (ICRA 2021), so I'm not a neutral party. Everything below is reproducible from the repo.

CABiNet is a dual-branch CNN for real-time semantic segmentation: a high-res spatial branch, a lightweight context branch (global aggregation + local distribution) over a MobileNetV3 backbone, fused with a small FFM. Published 2021, then it went quiet.

Read the original at Machine Learning