There's a quiet inefficiency hiding in plain sight on NVIDIA's RTX GPUs, and it's not the kind of thing most users will stumble upon by accident. A recent deep-dive into cuBLAS shows that batched FP32 operations on the RTX 5090 are only using about 40% of the available compute. That's not a rounding error or a niche edge case. It's the default behavior for every batched workload tested, from 256×256 matrices all the way up to 8192×8192 with a batch size of eight. The author of the analysis wrote a simple TMA double-buffer kernel that beats cuBLAS by 46% to 65% on the same hardware. That's not a marginal improvement. That's a fundamental gap between what the library ships and what the silicon can actually do.
What stands out is that this isn't a universal problem with cuBLAS. On NVIDIA's Pro 6000 and H200 GPUs, the library escalates through multiple tile sizes and reaches 73% to 82% FMA utilization. The RTX line just gets the short end of the stick. The kernel selection logic appears to skip the better paths entirely, leaving performance on the table for anyone doing batched FP32 work on consumer and enthusiast cards. The practical implication is direct: if you're on an RTX GPU and your workload involves batched matrix multiplies, you are likely leaving more than half of your card's compute idle. The fix isn't some exotic new hardware or a radical rewrite of your code. It's a matter of choosing a different kernel, or writing one yourself, which is exactly what the analysis demonstrates with a relatively straightforward implementation.
This also raises a broader point about how we trust default libraries. cuBLAS is treated as the gold standard for dense linear algebra, and for good reason, it's been battle-tested across decades and industries. But "battle-tested" doesn't mean "optimized for every SKU." The fact that a single developer's hand-written kernel can outperform the library by a wide margin on a flagship consumer GPU suggests that NVIDIA's optimization efforts are concentrated elsewhere. That's not a conspiracy. It's a resource allocation decision. The RTX line is massive, but it's also consumer hardware, and the company's enterprise customers likely get more attention in the tuning department. Still, when the gap is this large and the fix is this accessible, it's worth questioning whether the defaults we rely on are actually serving us, or just serving the path of least resistance.
The takeaway isn't that cuBLAS is broken or that NVIDIA is deliberately shortchanging RTX users. It's that performance isn't guaranteed by using a popular library. It has to be measured, and sometimes the measurement reveals something uncomfortable. The kernel isn't a miracle; it's a clear, well-documented example of what happens when you look under the hood and verify assumptions. For developers building on top of batched FP32 operations, the message is simple: run your own benchmarks, inspect the kernel selection, and don't assume the library knows your hardware better than you do. Because right now, on an RTX 5090, it doesn't. And the gap is too large to ignore.