1 min readfrom Machine Learning

[D] Howcome Muon is only being used for Transformers?

Our take

Muon has gained significant traction in the realm of large language model (LLM) training, yet its application beyond Transformers remains largely unexplored. Despite its announcement highlighting a new training speed record for Cifar-10, searches for Muon in convolutional networks yield minimal results. This raises important questions about its scalability and effectiveness in other contexts. If faster training typically correlates with improved model performance, what factors are limiting Muon's broader adoption? Have key papers on its applications been overlooked? Let's delve into this intriguing topic.

Muon has quickly been adopted in LLM training, yet we don't see it being talked about in other contexts. Searches for Muon on ConvNets turn up basically no results, despite its announcement including a new training speed record for Cifar-10. In my experience faster training usually comes with better final models, so what's the deal? Does it not actually scale? Have I missed papers?

submitted by /u/lukeiy
[link] [comments]

Read on the original site

Open the publisher's page for the full experience

View original article