There's a quiet assumption in the AI infrastructure world that Triton is for prototyping and CUDA is for production. This work challenges that assumption in a way that deserves attention. The author built a fused MoE dispatch kernel in pure Triton that outperforms Stanford's Megablocks at inference-relevant batch sizes, delivering 131% of the performance at 32 tokens and 124% at 128 tokens on Mixtral-8x7B with an A100. That's not a marginal gain. It's a direct challenge to the idea that you need hand-tuned, vendor-specific code to get serious performance out of modern model architectures.
What makes this notable isn't just the speedup, but the approach. The kernel fuses the gate and up projections so both GEMMs share the same input tile load, with SiLU computed directly in registers. That eliminates roughly 470MB of intermediate buffers per forward pass, a 35% reduction in memory traffic. The second contribution is a block-scheduled grouped GEMM that precomputes a mapping from block IDs to expert IDs and offsets, handling variable-sized expert batches in a single kernel launch without padding. These are not exotic tricks. They are practical, deliberate engineering decisions that reduce memory movement and kernel launch overhead, two of the biggest bottlenecks in inference today.
For developers working with Mixture-of-Experts models, this is a meaningful signal. It suggests that the performance ceiling for Triton is higher than many assumed, and that portability across hardware, the test suite also passes on AMD MI300X with zero code changes, does not have to come at the cost of inference efficiency. That's a tradeoff most would have called impossible a year ago. The fact that Megablocks pulls ahead at larger batches is expected, and it's honest reporting, but it doesn't diminish the core finding: for the batch sizes that matter in real-time inference, a pure Triton kernel can hold its own against hand-tuned CUDA.
The practical takeaway is straightforward. If you're building inference pipelines for MoE models, you should be exploring Triton as a first-class option, not a fallback. The code is open source, the writeup is detailed, and the results are reproducible. This isn't a claim about what might be possible in the future. It's a working implementation that already delivers measurable gains in the scenarios that matter most. Go read the writeup, run the tests, and see what happens when you stop treating CUDA as the only serious path forward.