**Our Take: Rose Rewrites the Rules of What an Optimizer Can Be**
This is not another optimizer wearing a fresh coat of paint. Matthew K. has released Rose, and it deserves serious attention from anyone who has felt the memory ceiling in PyTorch training. The breakthrough is painfully simple: a stateless optimizer that uses zero persistent memory for its state. That is not a marginal improvement. It means you can train larger models on the same hardware, or run experiments that previously required a cluster on a single GPU. The MNIST benchmarks speak for themselves, Rose achieves a top accuracy of 99.34% by epoch 11, while AdamW peaks at 99.30% in epoch 14. More importantly, Rose does this with a memory overhead of 0x versus AdamW's 2x. That is not hype. That is a measurable difference in what your workstation can do.
The practical implications are immediate. If you have ever had to downsample your batch size or reduce model dimensions to fit AdamW's twin copies of momentum and variance, Rose removes that constraint. The OpenAI parameter-golf test demonstrates a real-world win: Rose achieves a validation bpb of 2.2169, beating Adam's 2.2450, while using less memory. The developer is honest about benchmarks, yes, training loss can be higher while validation loss is lower, and the final output is what matters. That intellectual humility is refreshing, but the numbers are clear: Rose generalizes better on a smaller memory budget. For practitioners, this means you can iterate faster without chasing hardware upgrades.
We also note the human story here. Rose is named after the developer's mother, who loved hearing about his AI discoveries. That personal connection matters because it signals a tool built with care rather than corporate velocity. The Apache 2.0 license invites the community to test, break, and improve it. The developer explicitly asks for feedback, good or bad, which tells us this is a living project, not a polished debut meant for a press release. The visual comparison on Stable Diffusion training (AdamW left, Rose right) further validates that the gains hold beyond toy datasets.
Our final point is concrete: go run your own comparison. Clone the repo, swap out your optimizer import, and measure memory usage before and after. You will likely find that Rose allows you to train a model you previously could not fit, or speed up a pipeline that was bottlenecked by optimizer state. That is the only test that matters, and Rose invites it openly. This is not a replacement for every use case, but it is a genuine alternative for anyone constrained by memory or looking for better generalization. The proof is not in the claims, it is in your next training run.