Unlocking near-native FP8 performance with CUDA and PTX design

In his latest blog post, Daniel Vega-Myhre from Meta/PyTorch presents an insightful exploration of the MXFP8 GEMM design, achieving up to 99% of cuBLAS performance with CUDA and PTX.

2 min readMachine Learning

Daniel Vega-Myhre's deep-dive into unlocking near-native FP8 performance is exactly the kind of practical engineering work that moves the field forward. His walkthrough of GEMM design for MXFP8, covering every constraint and design challenge, doesn't just show how to make FP8 sing on CUDA and PTX; it gives practitioners a map they can follow. That matters because FP8 is the format everyone wants to use for large-scale training, but the gap between theoretical speed and real-world throughput has been stubbornly wide.

What Vega-Myhre demonstrates is that closing that gap requires understanding where the hardware actually lives, not where the spec sheet says it lives. His blog post methodically unpacks the trade-offs: memory alignment, shared memory occupancy, instruction scheduling, and the peculiarities of MXFP8's blocked exponent scheme. These aren't abstract concerns, they are the difference between a kernel that runs at 60% of peak and one that hits 90%. For teams building on Hopper-class GPUs, especially those working with models like DeepSeek-V3, this is the difference between a training run that finishes on schedule and one that doesn't.

The additional context from the PyTorch blog, where MXFP8 and DeepEP combine for up to 41% faster pre-training on B200 hardware, reinforces the same point: the hard work is in the details. Vega-Myhre's approach treats the GPU not as a black box but as a system with specific, navigable constraints. That is the mindset that turns a promising specification into a deliverable speedup. It is also the mindset that separates blog posts that teach from blog posts that merely announce.

If you are writing kernels for large-scale training, read this post with a notebook open. Vega-Myhre has done the debugging and the profiling so you don't have to start from scratch. The real takeaway is not that FP8 performance is possible, we already knew that, but that it is accessible to anyone willing to work through the hardware's actual behavior. That is an invitation worth accepting.

From Machine Learning

New blog post by Daniel Vega-Myhre (Meta/PyTorch) illustrating GEMM design for FP8, including deep-dives into all the constraints and design challenges introduced by MXFP8.

Link: https://danielvegamyhre.github.io/2026/03/29/mxfp8-gemm.html Original Tweet: https://x.com/vega_myhre/status/2038293614204445039

Read the original at Machine Learning