Tauon

Tauon brings faster training and lower loss to AI optimization

A new optimizer called Tauon is turning heads in the r/MachineLearning community, and for good reason.

4 min readMachine Learning
Tauon brings faster training and lower loss to AI optimization
Tauon: A new optimizer outperforming Muon on GPT-Mini (lower loss, ~8.5% faster step time) [P]
From Machine Learning

I’ve been working on a new optimizer called Tauon (turns out there is already "teon" but well if you have better idea, - i will gladly accept it! Anyway the core idea of optimizer is about polynomials and orthogonalization just like muon, the whole difference is that i managed to lower total number of steps (first through spectral filtering down to 3 steps then through coeff scheduling down to 2) + reduced matrix size (through dct-2). And I wanted to share some initial benchmark results...

Benchmark Setup: Trained a GPT-Mini (d_model=512, 6 Layers) on TinyShakespeare against Muon and AdamW.

Read the original at Machine Learning