KV Cache

Exploring how KV cache transforms LLM inference into an interactive runtime

Most LLM agents feel reactive because they move token by token, pausing for the model to catch up.

3 min readMachine Learning

Somewhere between the model weights and the harness that steers them, there is a layer of computing we tend to ignore. The Yandex research team is arguing that this layer, specifically the KV cache that carries an LLM's inference state, deserves attention as a runtime in its own right. Their post describes modifying this state directly to make models more interactive, building on earlier work like Hogwild! Inference and AsyncReasoning. The preview is compelling: a Qwen3.8-27B agent playing DOOM in real time, not by waiting for a full generation cycle, but by nudging the inference state as events unfold. This is a different kind of thinking about agent performance, and it is worth taking seriously.

Most discussions about improving LLM agents focus on two levers: the model itself or the harness that wraps it. The model is expensive to change, and the harness operates at a level of abstraction that can feel disconnected from the raw mechanics of generation. The team's question is whether the runtime, the internal state that determines how tokens are predicted, is an under-explored middle ground. It is a fair challenge. For anyone who has worked with Unlock LLM Training: A Practical Guide to Distributed Algorithms, the idea that inference is not just a black box but a system with its own state and bottlenecks is familiar. The KV cache is where the model's context lives, and treating it as something you can modify, rather than just read from, opens up new ways to think about responsiveness.

This is not about squeezing out a few extra tokens per second. The deeper implication is that interactivity, the ability for a model to react to its environment in near real time, might be a property of the runtime rather than the model or the prompt. The DOOM demo is a useful image here: a game environment requires continuous input, constant adjustment, and a sense of presence that batch-processed generations simply cannot provide. The team is not claiming to have solved general agent intelligence, but they are pointing at a specific, practical mechanism for making models feel more alive. For readers who have wrestled with Exploring Paragraph Structure: How LLMs Navigate Token Space, the connection is clear: if you can manipulate the internal coordinates of generation, you can change behavior without retraining a single weight.

Our take is that this is the right problem to be working on, and the question they pose deserves more attention. The harness is too abstract, and the model is too costly, but the runtime sits in a sweet spot where small changes can have outsized effects. What we would tell a reader who asks about this is simple: watch this space, because the ability to modify inference state is a lever that most of us have not even considered pulling. The concrete point to watch is whether this technique scales beyond a single agent playing a retro game. If it does, the next generation of LLM applications might not be built on better prompts or bigger models, but on a more intimate understanding of what happens inside the machine while it thinks.

From Machine Learning

Our research team has been exploring an alternative approach to achieving interactivity and better responsiveness with LLM systems.

One of the team members wrote up a post about it: https://research.yandex.com/blog/the-kv-cache-as-an-agent-runtime

Read the original at Machine Learning