SkewAdam

Tiered optimizer cuts MoE training memory demands by 97 percent

SkewAdam tackles the biggest pain point in MoE training head-on: optimizer state memory.

4 min readMachine Learning
Tiered optimizer cuts MoE training memory demands by 97 percent
SkewAdam: A tiered optimizer that cuts MoE state memory by 97% (fits a 6.7B MoE on a 40GB GPU) [R]

The math behind training large Mixture-of-Experts models has always felt like a quiet admission of defeat. You build a model that is supposed to be efficient, only to discover that the optimizer state, not the model, is what fills your GPU. SkewAdam, a new tiered optimizer from a team led by nuemaan, takes direct aim at this problem. The paper reports that AdamW spends 50.6 GB of state memory to update a 12.6 GB model. That is not a small overhead; it is the entire budget for most researchers. By rethinking how precision is allocated across parameters, SkewAdam drops that to 1.29 GB, a 97.4% reduction, and brings peak training memory from 81.4 GB down to 31.3 GB. That is the difference between needing a server and working with what you have.

The insight is refreshingly simple, and it is worth pausing on because it points to a deeper truth about how we treat optimization. SkewAdam does not treat every parameter as if it deserved the same level of care. The backbone, which accounts for 5% of parameters, keeps momentum and a factored second moment. The experts, which make up 95% of the model, rely on a factored second moment alone. The router, a sliver at under 0.01%, keeps an exact second moment. This is not a hack; it is a recognition that not all parameters are created equal. We have seen this kind of thinking elsewhere, like in Microsoft Open-Sources TauGrid to Simplify AI Workload Management on Kubernetes, where the focus is on smarter resource allocation rather than brute force. The same principle applies here: understand what needs precision and what can live with less.

What makes this particularly relevant is how it connects to the broader memory crunch we keep writing about. We have covered how The KV Cache Tax: Why Inference Servers Run Out of Memory Before Compute and how even the best hardware struggles under the weight of hidden costs. SkewAdam is the training-side answer to that same problem. It is not claiming to make models smarter; it is claiming to make them fit. For anyone who has had to rent a multi-GPU instance just to test a hypothesis, that is a tangible shift. The fact that a 6.78B MoE can now fit on a single 40GB GPU without sacrificing convergence or router stability is not a minor detail. It is the kind of practical win that opens the door for more researchers to experiment, rather than just those with access to large clusters.

Our take is straightforward: this is the right direction, and we would tell anyone asking about it to pay attention to the trade-offs. The paper does not claim to have invented a new optimization algorithm that beats Adam on quality. It has built something smarter about memory. That is a distinction worth holding onto. The real question is whether the tiered allocation holds up across different model sizes and training regimes. We would look at how the factored second moment performs on the experts over longer training runs, because that is where instability often creeps in. The router keeping an exact second moment is a smart hedge, but it is still a bet. The takeaway to quote is this: SkewAdam does not make MoE training easier by changing the model; it makes it easier by changing what we ask the hardware to remember. That is a trade-off we are happy to see explored.

From Machine Learning

Paper:https://arxiv.org/abs/2607.19058 Code (GitHub):https://github.com/nuemaan/skewadam

Hi everyone, I just published a preprint on a new optimizer designed to tackle the massive VRAM bottleneck in Mixture-of-Experts (MoE) training.

Read the original at Machine Learning