Adam

Why Adam Misses What Gradient Descent Naturally Preserves

The loss landscape of factored models hides a subtle trap, and this work exposes it with refreshing clarity.

4 min readMachine Learning
Why Adam Misses What Gradient Descent Naturally Preserves
The Loss Does Not See the Basis, But Adam Does [R]

The loss surface doesn't see the basis. Gradient descent doesn't either. Adam, it turns out, very much does. That is the central claim of a new paper that evaluates nine update rules on underdetermined matrix sensing, and the finding is as clean as it is provocative: optimizers that respect rotational invariance preserve GD's implicit low-rank bias, while those with per-coordinate second moments quietly throw that structure away. The authors test this by comparing every method at matched training loss, so no one can claim the winners simply underfit. The results split into two clusters. GD, shared-scalar Adam, Muon, and Shampoo keep the bias. Adam, RMSProp, Lion, signum, and Adafactor lose it.

What makes this worth your attention is not the binary result but the mechanism. The paper isolates the cause by taking Adam's denominator and gradually replacing its per-coordinate values with a single shared scalar. Recovery improves monotonically along that transition. That is a direct, causal demonstration that anisotropy, not adaptivity, is what breaks the structure. It is one thing to say per-coordinate scaling is theoretically unjustified in factored models. It is another to watch the degradation reverse in a smooth sweep. The same logic applies to the Muon optimizer, which behaves unexpectedly: exact on truly low-rank targets, but it cedes to GD at a crossover near 4% spectral tail energy. That reconciles two contradictory strands in the recent literature, where some report a strong spectral simplicity bias and others find it fits spurious features. Both are true, just at different points along the same axis.

The most actionable insight here is personal. The author applied the criterion to their own earlier optimizer and found that its per-coordinate clip was breaking the very structure it was designed to inject. Switching to a global norm clip improved recovery error from 0.347 to 0.220. That is a 37% relative improvement from a single, well-motivated change. If you are building or tuning optimizers for low-rank or overparameterized settings, this is the detail to steal. The caveats matter too. The headline 43-44% held-out error reduction on hyperspectral data relies on a train-only learning rate rule, and that rule assigns Adam the worst rate on its own grid. When each method picks its own optimal rate, the gap narrows considerably. The authors are honest about this, and they are right to keep the train-only rule because selecting on held-out data would inject the exact bias the experiment aims to avoid. The theoretical guarantees also cover memoryless rules only; momentum remains an empirical phenomenon.

When readers ask us what to make of this, we point to the deeper principle. The loss is a function of the product \(W = UV^T\), not of the individual factors. Any optimizer that treats coordinates as independent entities is making an implicit assumption about the basis, and that assumption has consequences. The practical takeaway is not that Adam is useless. It is that adaptivity and structure are in tension, and you need to know which one you are optimizing for. Watch the momentum question next. If a memoryless proof is the floor, finding the right way to incorporate momentum without reintroducing basis-dependence would be the ceiling. That is the open problem we would bet on.

From Machine Learning

In a factored model W = UV^T, the loss is invariant to rotations (U,V) → (UQ, VQ). Gradient Descent (GD) respects this property. Adam's per-coordinate second moment does not, because it depends on the specific basis in which the factors are written.

The claim is that this single property dictates whether optimizers retain or lose GD's implicit low-rank bias.

Read the original at Machine Learning