There is a quiet assumption hiding inside the current AI coding boom: that more context is always better. Feed the model your entire repository, every ticket, every stray design doc, and surely it will write better code. The presentation from Baruch Sadogursky and Patrick Debois challenges that premise head-on, and the timing could not be more critical. They argue that a bloated context window is not a feature; it is a performance drag. The insight that the right 300 tokens outperform 100,000 noisy ones is the kind of counterintuitive truth that separates teams who ship from teams who just burn tokens. We have been so busy scaling context that we forgot to ask whether the information was actually relevant. This is a welcome correction for anyone who has watched a coding agent confidently hallucinate its way through a monorepo. The conversation connects directly to Exploring Paragraph Structure: How LLMs Navigate Token Space, where token index is treated as a coordinate system; if we are going to navigate that space, we need better maps, not just bigger ones.
The practical fixes Debois and Sadogursky propose are refreshingly unglamorous. Lazy-loaded skills, versioned context artifacts, and externalized memory banks are not flashy demos. They are engineering discipline applied to a medium that often feels like magic. This is the right instinct. Instead of treating the model as an oracle that can absorb an entire codebase, they treat it as a capable assistant that needs a clean, curated workspace. The idea of versioning context artifacts is particularly sharp. Code is already versioned; why should the context we feed to an agent be any less rigorous? When a prompt fails, you need to know what changed. Treating context as a first-class artifact, with its own history and review process, turns debugging a bad output from guesswork into an investigation. For teams using Unlock ChatGPT for Work: A Practical Guide to Getting Started, this is the difference between a toy and a tool. You cannot build reliable workflows on top of a black box that eats everything you give it.
The most provocative piece here is the use of LLM-as-a-judge for evaluations. It is a step away from the assumption that human review is the only safety net. We are increasingly comfortable letting models critique other models, and this presentation suggests that for context engineering, that is not just acceptable; it is necessary. The speed at which prompts change demands an automated way to measure whether a tweak helped or hurt. This is not about removing humans from the loop. It is about making the loop fast enough to be useful. The takeaway worth quoting: "Context is a dependency, and it deserves the same discipline as any other dependency." That is the lens that turns raw markdown files into reliable agentic workflows. The open question we are left with is whether the broader industry will adopt this rigor, or if we will continue to be seduced by the sheer size of the context window. Watch for teams that start treating their prompt files like code reviews. Those are the ones who will build agents that actually last.
