LLM

When Milliseconds Matter: Teaching LLMs to Forget on Purpose

Most LLM runtimes treat memory as an endless resource, but this one treats it as a deadline.

3 min readTowards Data Science
When Milliseconds Matter: Teaching LLMs to Forget on Purpose

Most LLM inference runtimes treat memory as an afterthought. They cache, they evict, and they hope the model's next token arrives before the world moves on. A sharper question emerges: what happens when a physical deadline is non-negotiable? The answer, it turns out, is a system that refuses admission rather than miss a 33-millisecond robot control cycle. That is not a tweak. That is a different philosophy of computation, one where the model's internal state is managed with the same urgency as a real-time operating system manages interrupts.

The technical choices here are telling. Evicting KV cache by meaning instead of age is a quiet rebellion against the lazy assumption that recency equals relevance. In a control loop, a stale token is not just unhelpful; it is dangerous. And writing the entire runtime in hand-written CUDA, with no cuBLAS and no libtorch, signals a willingness to abandon convenience for control. This is not about being clever for its own sake. It is about recognizing that the biggest bottleneck in AI is rarely the model itself. It is the plumbing around it. We have seen similar tensions in other corners of the field, like when Unlock LLM Training: A Practical Guide to Distributed Algorithms unpacks how communication overhead often dwarfs compute time. The lesson is the same: the system is the model.

But here is where we push back. This is framed as a technical achievement, and it is. Yet the deeper implication is about trust. If an LLM is going to control a robot, it must be able to say "no" when it cannot keep up. That is a form of honesty we rarely ask of our AI systems. We prefer models that hallucinate confidence over ones that admit latency pressure. This is why we would tell a reader to pay attention to the admission control mechanism, not just the CUDA kernels. A model that refuses a request because it cannot guarantee a deadline is a model that understands its own limits. That is a feature, not a bug. It also invites a broader conversation about verification, one that echoes the concerns raised in Verify Your AI's Understanding: A Simple Check for Tax Season, where the real risk is not a wrong answer but an unexamined one.

The takeaway we hope sticks is this: forget the right things, but only if you also remember what you are for. In a robot control loop, forgetting the wrong token is a crash. Forgetting the right token is a save. The system described here is not just optimizing for accuracy; it is optimizing for survival. That is a mindset shift. We would tell a curious engineer to stop asking "how fast is the model?" and start asking "what is the model allowed to miss?" Because in the end, the future of AI is not about doing more. It is about knowing what to leave out, and having the courage to refuse the rest. Watch for the next wave of runtimes that treat refusal as a first-class operation, not an error state. That is where the real progress will be made.

From Towards Data Science

Most LLM inference runtimes have no idea a physical deadline exists. This one refuses admission rather than miss a 33ms robot control cycle, evicts KV cache by meaning instead of age, and is written entirely in hand-written CUDA — no cuBLAS, no libtorch.

The post Can an LLM Forget the Right Things? appeared first on Towards Data Science.

Read the original at Towards Data Science