Machine Learning

Smarter retrieval reduces inputs by 4-5x while improving chat accuracy

Automatic budget selection in chat retrieval is a puzzle many of us know well, and this user's experience at 25% accuracy is telling.

4 min readMachine Learning
Smarter retrieval reduces inputs by 4-5x while improving chat accuracy
Input 4-5x Reduction with sentence and keyword based trie on chat. [P]

The request here isn't for a generic summary; it's for a point of view on a very specific technical pain point. The original poster is wrestling with a real problem: their trie-based retrieval system is hitting 25% budget and matching benchmark accuracy, sometimes even outperforming on live chat, yet it still retrieves too much. That last part is the tell. The system is close, but it lacks the final layer of discernment that separates a useful tool from a merely accurate one. The architecture isn't a failure; it's a signal that it's ready for the next step.

This is where the conversation gets interesting for anyone building on large language models. You're not looking for a fundamentally new system; you're looking for a smarter selector. The poster mentions CELF as the current method for automatic budget selection, but the honest take is that CELF, while solid, is a blunt instrument for this context. It optimizes for a global score, not for the nuanced reality of a chat where the user's intent shifts from question to statement to command within a single exchange. The real opportunity here is to move from a static budget to a dynamic one, where the trie's own confidence scores, sentence boundaries, and keyword density feed into a second-stage model that decides *how much* context is truly needed. That's not just an incremental improvement; it's the difference between a retrieval system that works and one that feels intuitive.

For our readers, the practical takeaway is to stop treating retrieval as a single optimization problem. The fact that 25% matches benchmarks is a milestone, not a destination. It means the trie is doing its job on precision; the gap is in recall management. If you're building something similar, look at your own threshold as a starting point, not a rule. Test where the marginal return on additional tokens flatlines. The poster's experience suggests that the sweet spot is dynamic, and the next step is to build a lightweight classifier that predicts the optimal budget per query, rather than relying on a fixed percentage. That's the kind of work that turns a good demo into a production-ready feature.

The specific detail to watch is the behavior on actual chat input versus benchmarks. If the trie is performing *better* on real data, that's a sign your benchmark is too easy or your real distribution has more redundancy to exploit. Either way, the open question is whether you can build a feedback loop that learns from that gap. We'd tell anyone in this position to stop chasing a better CELF variant and start experimenting with a hybrid approach: keep the trie for candidate generation, but add a second-pass ranker that prunes based on sentence-level relevance. That's where the 4-5x reduction becomes a consistent reality, not a hopeful benchmark. For a deeper dive into how token-level decisions shape this kind of architecture, our piece on Exploring Paragraph Structure: How LLMs Navigate Token Space is a good companion. And if you're thinking about scaling this beyond a single session, the fundamentals in Unlock LLM Training: A Practical Guide to Distributed Algorithms apply to the inference side too. Finally, for those applying this to user-facing tools, the setup principles in Unlock ChatGPT for Work: A Practical Guide to Getting Started are directly relevant. The next step is clear: build the selector, watch the live data, and let the budget follow the query, not the other way around.

From Machine Learning

Currently struggling with an automatic budget selection, at 25% it’s very similar to benchmarks accuracy and seems even better on actual chat input however it many times retrieves too much. It would be nice to add an algorithm that actually can determine better retrieval other then CELF.

Read the original at Machine Learning