KV Cache

Explore the Hidden Geometry Inside Your Model's Working Memory

The KV cache isn't a flat list; it's a navigable vector space, and this researcher turned that observation into a working system.

4 min readMachine Learning

The KV cache has always been the quiet workhorse of inference, the thing we optimize around without always interrogating what it actually is. This post reframes it in a way that should feel obvious in hindsight: the cache is not a flat array of past tokens, it is a learned geometry. The keys encode relationships, the queries navigate them, and attention is just a similarity search that happens to run exhaustively. That framing changes the engineering problem from "how do I store more?" to "how do I route to the right neighborhood?" A measured result, a 16 to 31 times reduction in physical KV reads on a frozen Qwen3.5-2B at 32k context while still retrieving a planted long-range needle, is the kind of concrete number that makes the conceptual shift worth taking seriously.

This is where the idea stops being academic and becomes practical. If relevance is not uniformly distributed across the cache, and queries tend to concentrate on small regions of old context, then exhaustive scanning is not just inefficient, it is architecturally lazy. This is essentially proposing a hierarchical index for working memory, one that mirrors what information retrieval systems have done for decades, but applied to the internal state of a transformer. The fact that window-only and random-routing controls collapse is the right kind of control experiment. It tells us the geometry is doing real work, not just that the model happens to be robust to noise. For anyone building on top of open-weight models, this is a direct path to reducing latency and memory bandwidth without retraining or fine-tuning. It is also a reminder that the Unlock LLM Training: A Practical Guide to Distributed Algorithms conversation often focuses on the forward pass and weight updates, but inference-time memory management is where the real user-facing gains live.

What we find compelling is the honesty in the framing. The author admits the initial post was poorly pitched, that they had already built and measured the mechanism before asking the conceptual question. That is not a flaw, it is the right way to do research. Too often we see ideas that are all vision and no validation. Here, the vision is backed by a minimal runnable demo, which is exactly the kind of artifact that lets others verify and build. This is also a useful lens for thinking about the broader trajectory of the field. The recent NeurIPS Acceptance Raises Questions About AI Review Justifications highlights how much weight we put on formal evaluation, but practical, reproducible results like this one often matter more than a polished paper. And as the Explore the Future of AI Deployment: Key Topics at QCon AI New York agenda suggests, production systems are where these ideas will live or die.

Our take is straightforward: stop treating the KV cache as a memory problem and start treating it as a navigation problem. The open question is how far this generalizes. The result here is on a specific model at a specific context length, and the routing mechanism is still relatively simple. But the principle, that attention is search and search can be indexed, is likely to hold across scales. The detail to watch is whether this holds up when the cache grows beyond a single document or becomes multi-tenant. If it does, we will look back at exhaustive attention the way we now look at full table scans in databases, as a correct but primitive baseline.

From Machine Learning

I've been doing some research on this question:

At inference time a large part of a model's working memory lives in the KV cache, plus whatever external memory the harness bolts on. I've been poking at the storage-and-retrieval side of this, treating that cache as an index, and what stands out is that it isn't a flat list. It's a structured set of vectors with a navigable geometry, since the keys carry the model's learned sense of what relates to what.

Read the original at Machine Learning