1 min readfrom Machine Learning

[D] MXFP8 GEMM: Up to 99% of cuBLAS performance using CUDA + PTX

Our take

In his latest blog post, Daniel Vega-Myhre from Meta/PyTorch presents an insightful exploration of the MXFP8 GEMM design, achieving up to 99% of cuBLAS performance with CUDA and PTX. This comprehensive piece delves into the constraints and design challenges associated with FP8, offering valuable perspectives for developers and researchers alike. Vega-Myhre’s work highlights the innovative approaches that enhance performance and efficiency in data processing. Discover the full analysis and its implications for future advancements in machine learning at the provided link.

New blog post by Daniel Vega-Myhre (Meta/PyTorch) illustrating GEMM design for FP8, including deep-dives into all the constraints and design challenges introduced by MXFP8.

Link: https://danielvegamyhre.github.io/2026/03/29/mxfp8-gemm.html
Original Tweet: https://x.com/vega_myhre/status/2038293614204445039

Additional resources:
MXFP8 and DeepEP for DeepSeek-V3 on B200 w/ TorchTitan: https://pytorch.org/blog/enabling-up-to-41-faster-pre-training-mxfp8-and-deepep-for-deepseek-v3-on-b200-with-torchtitan/

submitted by /u/Benlus
[link] [comments]

Read on the original site

Open the publisher's page for the full experience

View original article