Beyond Transformers: Unlocking Muon's Potential for Every Model

Muon has gained significant traction in the realm of large language model (LLM) training, yet its application beyond Transformers remains largely unexplored.

2 min readMachine Learning

The silence around Muon outside large language model training is a missed opportunity, not a sign of failure. The optimizer that set a new speed record on CIFAR-10 with convolutional networks deserves far more attention than it is getting.

The community's narrow focus makes sense historically. Transformers dominate the conversation, and new optimizers often get tested where the hype is loudest. But the original Muon results on ConvNets were not a fluke. Faster training correlates strongly with better final models across many architectures, and a method that accelerates convergence on small-scale vision tasks should scale to larger ones. The fact that searches for Muon on ConvNets return almost nothing suggests a collective blind spot, not a limitation of the optimizer itself.

What this means for practitioners is straightforward. If you are training convolutional models for image classification, object detection, or any vision task, Muon is worth a serious look. The optimizer's design, leveraging matrix orthogonalization, is architecture-agnostic in principle. There is no fundamental reason it should work only on transformers. The gap in the literature is a gap in experimentation, not in theory. Early adopters who run controlled comparisons on their own vision workloads stand to gain a meaningful speed advantage while the rest of the field catches up.

The path forward is simple: test Muon on your next ConvNet training run. Run it against AdamW on the same architecture, same data, same schedule. If the speed record on CIFAR-10 holds in your setting, you have found a practical win. If it does not, publish the negative result so others learn. Either outcome moves the field forward more than silence does.

From Machine Learning

Muon has quickly been adopted in LLM training, yet we don't see it being talked about in other contexts. Searches for Muon on ConvNets turn up basically no results, despite its announcement including a new training speed record for Cifar-10. In my experience faster training usually comes with better final models, so what's the deal? Does it not actually scale? Have I missed papers?

Read the original at Machine Learning