The silence around Muon outside large language model training is a missed opportunity, not a sign of failure. The optimizer that set a new speed record on CIFAR-10 with convolutional networks deserves far more attention than it is getting.
The community's narrow focus makes sense historically. Transformers dominate the conversation, and new optimizers often get tested where the hype is loudest. But the original Muon results on ConvNets were not a fluke. Faster training correlates strongly with better final models across many architectures, and a method that accelerates convergence on small-scale vision tasks should scale to larger ones. The fact that searches for Muon on ConvNets return almost nothing suggests a collective blind spot, not a limitation of the optimizer itself.
What this means for practitioners is straightforward. If you are training convolutional models for image classification, object detection, or any vision task, Muon is worth a serious look. The optimizer's design, leveraging matrix orthogonalization, is architecture-agnostic in principle. There is no fundamental reason it should work only on transformers. The gap in the literature is a gap in experimentation, not in theory. Early adopters who run controlled comparisons on their own vision workloads stand to gain a meaningful speed advantage while the rest of the field catches up.
The path forward is simple: test Muon on your next ConvNet training run. Run it against AdamW on the same architecture, same data, same schedule. If the speed record on CIFAR-10 holds in your setting, you have found a practical win. If it does not, publish the negative result so others learn. Either outcome moves the field forward more than silence does.