7 min readfrom Machine Learning

CABiNet (ICRA 2021) vs YOLO26-sem on UAVid: accuracy, compute, and GPU latency [P]

Our take

Published in 2021, CABiNet (ICRA 2021) is a dual-branch CNN for real-time semantic segmentation that has now been revisited and benchmarked against YOLO26-sem on the UAVid dataset. Our controlled experiment, reproducible from the linked repository, reveals that CABiNet achieves a higher mIoU (67.14% vs 64.41%) with significantly lower GPU latency (4.44 ms vs 13.09 ms) than YOLO26x-sem. This demonstrates that a purpose-built, efficient architecture can outperform larger, multi-task models, particularly
CABiNet (ICRA 2021) vs YOLO26-sem on UAVid: accuracy, compute, and GPU latency [P]

The recent comparative analysis of CABiNet (ICRA 2021) and YOLO26-sem on the UAVid dataset, as detailed in a Reddit post, offers a fascinating glimpse into the evolving landscape of real-time semantic segmentation. The author, a primary contributor to CABiNet, meticulously outlines a controlled benchmark, aiming to fairly assess a purpose-built, efficient 2021 architecture against a more recent, general multi-task model. This is particularly relevant as the field rapidly progresses, with models like those explored in OpenAI Details GPT-Live’s Architecture for Continuous Stateful Voice Interaction demonstrating the increasing complexity and breadth of AI applications. The decision to standardize data representation, class weighting, and evaluation protocols while allowing each model to retain its unique training recipe is a clever approach to isolate architectural differences and provides valuable insights into the trade-offs between efficiency and accuracy. The resulting data, openly shared and demonstrated via a live demo, highlights the importance of nuanced performance analysis beyond simple headline metrics.

What makes this analysis particularly compelling is the author's candor about potential confounding factors, such as the asymmetric initialization – YOLO26-sem benefiting from pre-training on larger datasets like Cityscapes and ADE20K. This acknowledgement underscores the value of transparent and reproducible research, a principle championed in initiatives like the massive TikTok dataset release [I scraped 5.94 billion TikTok videos and 3.23 billion profiles in 3 weeks. Uploaded full dataset to Hugging Face for free. Step by step tutorial and code below. [P]](/post/i-scraped-5-94-billion-tiktok-videos-and-3-23-billion-profil-cmtkekd2i01nfrgeddu2u8ft5). The findings regarding MobileNetV3’s depthwise convolutions – being computationally efficient but not GPU-latency-cheap – are a crucial observation for developers optimizing for real-time performance, especially in resource-constrained environments like aerial robotics. The breakdown of performance differences across specific classes (Human, Static Car, Moving Car) provides granular detail that goes beyond aggregate metrics, demonstrating where each architecture excels and where improvements might be targeted. The release of the weights and configurations further enhances the accessibility and reproducibility of the research, aligning with the spirit of open-source innovation exemplified by projects like [We released TontaubeV1, a character-level TTS model for long-form generation [P]](/post/we-released-tontaubev1-a-character-level-tts-model-for-long-cmtk1qwzy01hfrgedjb9ypi2g), which prioritizes accessibility and community contribution.

The conclusion that CABiNet occupies the higher-accuracy end of the accuracy/latency Pareto frontier is significant. It suggests that purpose-built architectures, even those from earlier iterations, can still offer compelling performance advantages when optimized for specific tasks. The analysis doesn't claim universal superiority; rather, it demonstrates a clear trade-off and highlights the importance of selecting the right tool for the job. The author’s call for criticism regarding the standardization approach is particularly insightful, encouraging a broader discussion about how best to evaluate and compare models from diverse lineages, especially as AI models become increasingly complex and specialized. The measured performance on different datasets, readily available in the repository, provides a more complete picture of each model’s adaptability.

Looking ahead, it’s worth considering how these findings might inform the development of future aerial perception systems. Will we see a resurgence of specialized architectures like CABiNet, tailored for specific domains, or will the trend continue toward increasingly general, multi-task models like YOLO26? The efficiency gains offered by MobileNetV3, coupled with the accuracy improvements demonstrated by CABiNet, suggest a potential path towards achieving both real-time performance and high accuracy in aerial semantic segmentation. The question remains: how can we best leverage the strengths of both specialized and general models to create truly adaptable and robust AI systems for increasingly complex aerial environments?

CABiNet (ICRA 2021) vs YOLO26-sem on UAVid: accuracy, compute, and GPU latency [P]

Disclosure up front: I'm the original first author of CABiNet (ICRA 2021), so I'm not a neutral party. Everything below is reproducible from the repo.

Background

CABiNet is a dual-branch CNN for real-time semantic segmentation: a high-res spatial branch, a lightweight context branch (global aggregation + local distribution) over a MobileNetV3 backbone, fused with a small FFM. Published 2021, then it went quiet.

I came back this year, rebuilt the repo (PyTorch 2.x, Hydra, AMP, EMA, poly-LR, OHEM loss, CI + tests), and used it to ask one question on **UAVid**, the aerial dataset the original paper targeted: how does a purpose-built 2021 efficient architecture compare to a 2026 general multi-task model with a dedicated semantic-segmentation variant?

What's actually controlled (and what isn't)

Both models run off the same converted dataset and splits, the same ENet inverse-log class weighting (`cls_pw=0.5`), EMA weights for eval, and the same evaluation protocol: single-scale, no test-time augmentation. What is not matched:

| Axis | CABiNet | YOLO26-sem | Potential advantage | | --- | --- | --- | --- | | Initialization | ImageNet-pretrained MobileNetV3 backbone; seg layers random | full net pretrained on Cityscapes + ADE20K | potentially favors YOLO | | Epoch budget | 5000 (early stop, patience 100) | 500 (early stop, patience 50) | potentially favors CABiNet | | Optimizer / schedule | SGD + poly decay, decoder LR ×10 | SGD + cosine | different | | Loss | OHEM-CE + aux deep supervision | CE + Dice + aux | different | | Extra augmentation | none | mosaic 0.8, copy-paste 0.15 | potentially favors YOLO | 

So this is not an architecture-only ablation. It's a controlled benchmark: the data representation, class weighting and evaluation are standardized, while each model keeps a model-specific training recipe. None of the rows above is an isolated experiment, so I haven't measured how much any single one is worth.

Results — UAVid test split, 1024×1024, single-scale

| Model | mIoU (%) | Params (M) | FLOPs (G) | FP16 latency* | FP16 FPS | | --- | --- | --- | --- | --- | --- | | **CABiNet (MobileNetV3-L)** | **67.14** | 9.17 | 54.8 | 4.44 ms | 225 | | **CABiNet (MobileNetV3-S)** | 65.25 | 5.36 | 44.1 | 3.09 ms | 324 | | YOLO26x-sem | 64.41 | 40.16 | 430.9 | 13.09 ms | 76 | | YOLO26l-sem | 63.28 | 17.87 | 192.4 | 7.54 ms | 133 | | YOLO26m-sem | 61.98 | 14.32 | 152.3 | 5.71 ms | 175 | | YOLO26s-sem | 61.69 | 6.50 | 44.4 | 2.52 ms | 396 | | YOLO26n-sem | 58.17 | 1.63 | 11.4 | 2.23 ms | 449 | *\*RTX 4070 SUPER, batch 1, pure model forward pass (no pre/post), 200 iters after 30 warmup, measured by me. Params are architecture-only; FLOPs are analytic forward-pass at 1024² (thop for CABiNet, Ultralytics profiler for YOLO26; both report FLOPs = 2×MACs).* 

UAVid mIOU vs FP16 Latency

The dashed line is the accuracy/latency Pareto frontier: YOLO26n and YOLO26s sit on it as legitimate lower-latency points, while YOLO26m/l/x are dominated, each being both slower and less accurate than at least one CABiNet variant. CABiNet occupies the higher-accuracy end of the frontier.

Three things worth pulling out:

  1. Near-iso-compute: CABiNet-S vs YOLO26s. ~44 GFLOPs each (44.1 vs 44.4), CABiNet-S has slightly fewer params (5.36M vs 6.50M), and they're within 0.6 ms on this GPU, yet CABiNet-S is +3.6 mIoU (65.25 vs 61.69). YOLO26s is still the faster model, so this is a clean accuracy/latency trade, not a universal win.
  2. Higher-accuracy end: CABiNet-L vs YOLO26x. CABiNet-L is +2.7 mIoU and ~3× lower forward latency (4.44 vs 13.09 ms). It's not that CABiNet is the fastest model (YOLO26n/s are faster); it's that it reaches higher accuracy without moving into the latency/compute regime of YOLO26m/l/x.
  3. Not universally better. On VDD and AeroScapes (same matched eval), YOLO26 s-and-up pull ahead of CABiNet-Large, which lands mid-pack there. Numbers and configs in the repo.

MobileNetV3's depthwise convs are FLOP-cheap but not GPU-latency-cheap, which is why the frontier looks the way it does. The story is accuracy per millisecond at the higher-accuracy end, not "smallest and fastest."

Qualitative — CABiNet-L vs YOLO26x-sem

Where the +2.7 mIoU comes from. Per-class IoU on the UAVid test split, matched single-scale:

| Class | CABiNet-L | YOLO26x-sem | Δ | | --- | --- | --- | --- | | Human | 28.3 | 21.1 | **+7.2** | | Static Car | 57.2 | 51.3 | **+5.9** | | Moving Car | 71.9 | 66.8 | **+5.1** | | Tree | 80.3 | 78.2 | +2.1 | | Vegetation | 64.1 | 63.3 | +0.8 | | Road | 80.3 | 79.8 | +0.5 | | Clutter | 67.8 | 67.3 | +0.5 | | Building | 87.1 | 87.4 | −0.2 | 

UAVid Test Set Qualitative Comparison

The gap is almost entirely the small / thin classes: people and vehicles. On the big region classes the two are within half a point, and YOLO26x is marginally ahead on Building.

Two UAVid test frames, both single-scale; columns are input · YOLO26x-sem · CABiNet-L · ground truth. Row 2 shows a failure mode behind the Static-Car number: YOLO26x collapses the parking-lot structure into one Static-Car/Clutter mass and bleeds Building into the lot, while CABiNet-L tracks the ground truth more closely. These two frames were chosen to illustrate the per-class differences above, not as a representative random sample.

Scope / limitations

  • UAVid only (see point 3 above). The VDD / AeroScapes numbers and configs are in the repo; I'm leading with UAVid because that's where the result is clean, not hiding the rest.
  • Single training run per config: no seed sweep, no variance estimate. The observed ~2.7 mIoU CABiNet-L vs YOLO26x gap is large relative to the smaller differences in this table, but I haven't established statistical significance. I wouldn't over-read anything under ~1 point.
  • Latency is a clean-room forward pass on one consumer GPU. No TensorRT/ONNX, no Jetson, no full-frame sliding-window cost (UAVid source frames are 4K; CABiNet tiles, YOLO resizes, so end-to-end numbers would differ). Read these as model-level GPU measurements, not deployment throughput.
  • The initialization is asymmetric: YOLO26-sem starts from Cityscapes + ADE20K pretraining, CABiNet only from an ImageNet-pretrained backbone. This likely gives YOLO26 a transfer learning advantage on aerial data, though its magnitude isn't measured here. CABiNet reaching higher UAVid accuracy from the less domain-specific start is part of what makes the result interesting, but it stays a confound.

Open-sourced

Links

The criticism I'd most like: is standardizing the data representation, class weighting and evaluation, while letting each model keep its native training recipe, a useful way to compare architectures from different lineages? If not, what would you standardize or change instead?

submitted by /u/Naive-Explanation940
[link] [comments]

Read on the original site

Open the publisher's page for the full experience

View original article
CABiNet (ICRA 2021) vs YOLO26-sem on UAVid: accuracy, compute, and GPU latency [P] | Beyond Market Intelligence