There's a quiet revolution happening in long-context inference, and it's not coming from a massive lab with unlimited compute. It's coming from a single developer who decided that reproducibility matters as much as results. The open-source implementations of Cartridges and STILL bring neural KV-cache compaction down to a single GPU, and that's not just impressive, it's a shift in who gets to play with these ideas.
Let's be clear about what this means in practice. We've been told for a while that long-context windows are the future, but the memory cost of key-value caches has always made that future feel exclusive. You needed serious hardware just to keep a conversation going, let alone process entire books or codebases. What these repos do is show that you don't need that scale anymore. Cartridges compresses caches for specific corpora, while STILL takes a more reusable approach, compacting caches in a way that generalizes across contexts. And because both are open source, you can actually read the code, run the benchmarks, and see where the tradeoffs are. That's not a black box, that's a toolkit.
What's particularly refreshing is the motivation behind the work. The goal wasn't to publish a flashy blog post or claim a new state of the art. It was to make the ideas easy to inspect and run. That's a human-centered approach to AI research, and it's one we should get behind. Too often, papers and summaries leave you squinting at equations, wondering if the results are real or just carefully selected. Here, you have benchmark code and readable implementations. You can test these methods against full-context inference, truncation, and each other. That's the kind of transparency that builds trust, and it's also how progress actually happens, through iteration and scrutiny, not hype.
For developers and researchers, this is a practical invitation. If you've been holding off on exploring KV-cache compaction because it felt too complex or too expensive, this removes those barriers. You can load these repos, run them on a single GPU, and start understanding the systems tradeoffs yourself. That's not just convenient, it's empowering. It means the next breakthrough might not come from a well-funded team, but from someone who took the time to build on this work and share it. That's the kind of future we want to be part of, and it's already here. Start with the code.