search engine

How one developer slashed AI search costs while keeping queries accurate

Building a search engine that actually understands "a game like Hades that feels cozier and is co-op" is a hard problem.

4 min readMachine Learning

The developer behind indiedex.gg has done something more important than build a clever game search engine. They've demonstrated a practical pattern for slashing AI costs and latency without sacrificing quality, and that pattern is one our readers should explore immediately. By swapping a general-purpose DeepSeek extraction call for a structured, closed-question model (TypeSafe Jev on OpenRouter's Decisions API), they cut extract time by roughly 10×, reduced end-to-end route time nearly in half, and dropped cost about 6×, all while holding accuracy on their test set. This is the kind of engineering trade-off that deserves attention, because it sidesteps the expensive habit of treating every query like a free-form chat.

What makes this approach instructive is its focus on intent extraction rather than conversation. The indiedex.gg pipeline takes a compound query like "a game like Hades that feels cozier and is co-op," turns it into structured pieces (reference game, filters, tags, vibe), and then lets traditional tools handle the rest. That's a fundamentally different architecture than the chatbot-style completions many teams default to. We recently covered a similar theme in "Explore a smarter AI model that lifts complex work without the premium price," where OpenAI's GPT-6.1 Sol showed that focused improvements in professional tasks can deliver more value than broad capability boosts. And in "Cut token costs by routing text-heavy images away from pixel processing," another developer demonstrated that routing inputs to the cheapest model that can handle them is a reliable cost lever. The indiedex.gg case extends that logic: instead of routing to a cheaper model for the same open-ended task, it redesigned the task itself to be a set of closed questions, making it naturally cheaper and faster.

The practical takeaway for anyone building "natural language in → structured intent out" systems is clear: not every input needs a full generative model. When your goal is to extract a fixed set of fields, yes/no answers, categorical choices, numeric ranges, you can use a model designed for that exact purpose. The developer here retained their existing regex helpers and search pipeline, swapping only the extraction step. That modularity matters. It means you can upgrade or swap components without rewriting your entire stack. The quality metric worth watching is whether this holds at scale, across more diverse and ambiguous queries than the 50-case bakeoff tested. The developer reports high agreement on hard filters and title resolution, but real-world user queries will stretch those boundaries.

One specific detail to watch: the developer notes that Jev "doesn't write a free-form answer. It answers a fixed set of closed questions in parallel." That parallel execution is likely a significant contributor to the speed gains, and it suggests a design principle, parallelize your model calls wherever the output dimensions are independent. For teams building decision engines, smart filtering, or auto-moderation systems, this pattern offers a concrete starting point. The question now is how many other pipelines can be refactored the same way, and whether the ecosystem of structured-output models will grow to match the demand.

From Machine Learning

I built indiedex.gg, a hidden-gems game search engine on steam data. Search is the front door: type something like “a game like Hades that feels cozier and is co-op,” and it should actually understand that mix of reference + vibe + filters.

Those compound queries used to go through a DeepSeek extraction call that turned the sentence into structured pieces (reference game, filters, tags, vibe). It worked, but it was slow and expensive on the hot path.

Read the original at Machine Learning