•1 min read•from Machine Learning
[D] Howcome Muon is only being used for Transformers?
Our take
Muon has gained significant traction in the realm of large language model (LLM) training, yet its application beyond Transformers remains largely unexplored. Despite its announcement highlighting a new training speed record for Cifar-10, searches for Muon in convolutional networks yield minimal results. This raises important questions about its scalability and effectiveness in other contexts. If faster training typically correlates with improved model performance, what factors are limiting Muon's broader adoption? Have key papers on its applications been overlooked? Let's delve into this intriguing topic.
Muon has quickly been adopted in LLM training, yet we don't see it being talked about in other contexts. Searches for Muon on ConvNets turn up basically no results, despite its announcement including a new training speed record for Cifar-10. In my experience faster training usually comes with better final models, so what's the deal? Does it not actually scale? Have I missed papers?
[link] [comments]
Read on the original site
Open the publisher's page for the full experience