ALHR

Sparse attention model reads 94% fewer keys while keeping accuracy

Reading just 30 keys per query instead of 512 is a compelling proof point.

3 min readMachine Learning

The race to make attention mechanisms less wasteful just gained another serious contender. ALHR, Adaptive Learnable Hierarchical Routing, reads only about 30 keys per query on a 1024-token test, compared to 512 for a dense baseline. It keeps top-1 accuracy at 92.1% versus 94.9%. That is a 35.3x compression of the KV cache, and the peak VRAM profile shifts from quadratic to linear scaling. We should pay attention to this because it addresses the exact bottleneck that limits how far we can stretch context windows, a theme we have explored in Two new attention models slash redundancy to stretch context further.

The clever part is the mechanism itself. ALHR builds static binary trees and uses learnable functions to route queries through them, dropping branches that do not contain relevant keys. This is not a sparse approximation that sacrifices structure for speed. It is a hierarchical decision process, and the training phase does lean on a dense teacher, which is worth noting. The model learns to prune aggressively while still recovering most of the accuracy. That gap, 94.9% down to 92.1%, is a real tradeoff, but the efficiency gain is not marginal. It is a step change. The authors are honest about limitations: full-scale tests are pending, and training remains quadratic. Inference, however, drops to NlogN. That is where the practical value lives.

What does this mean for you, the person actually building with these models? It means the cost of long-context inference does not have to explode as your data grows. Dense attention reads every key for every query, which is why peak VRAM scales quadratically. ALHR reads a fraction of the keys and scales linearly at inference. For teams running large-scale retrieval or document analysis, that difference is the difference between a feasible deployment and a budget-breaking one. The compressed cache also means more tokens can stay resident in memory, which directly supports the kind of extended context work we have seen in How small AI models closed the gap on a test built for humans. Smaller models, smarter routing, and less redundancy are converging on the same outcome: more capability per unit of compute.

Our take is that ALHR is not just another pruning trick. It is a practical demonstration that hierarchical routing is a viable path forward for attention efficiency. The dense teacher during phase one is a dependency, not a flaw. It shows that we can distill routing behavior from expensive models into cheaper inference-time architectures. The open question is whether the 2.8% accuracy drop holds at larger scales. If it does, this approach deserves serious consideration for production systems. If the gap widens, the tradeoff becomes harder to justify. Watch the full-scale results. That is the detail that will determine whether ALHR becomes a standard tool or a promising footnote.

From Machine Learning

ALHR - Adaptive Learnable Hierarchical Routing, uses static binary trees and learnable functions to minimize the amount of keys to be read.

It does use a dense teacher while phase 1 of training however.

Read the original at Machine Learning