Stop routing every question to the AI before you filter the easy ones.

Most teams building RAG systems route every ambiguous case straight to the language model, trusting retrieved context to sort it out.

4 min readVentureBeat
Stop routing every question to the AI before you filter the easy ones.

The most expensive sentence in enterprise AI isn't the one that returns a wrong answer. It's the one that sounds right while doing it, especially when a compliance officer is the one reading it six months later. Vineet Vijay's piece on cascade architecture for RAG systems lands at a moment when too many teams are still treating the language model as the default first responder for every ambiguous case. That approach works in a demo because demos don't have auditors. Real systems do. And as Vijay points out, the bill for that architectural bet comes due in three forms: auditability, cost at scale, and the quiet inconsistency of LLMs on cases that should be deterministic. We'd go further: if your classification pipeline can't explain a decision without rerunning inference, you don't have a system yet. You have a slot machine with a confidence score.

The cascade model Vijay describes is not a clever workaround. It's the correct default for any high-stakes domain, and the 6x cost reduction is almost beside the point. The real win is that deterministic stages force you to encode your institutional knowledge into rules and retrievals, which means the system's reasoning is inspectable by design. When a case does reach the LLM, you've already done the hard work of isolating the genuinely ambiguous residue. That's where his asymmetric risk prompt becomes the quiet hero. Most teams default to neutral language like "assess whether to approve," which asks the model to treat a false negative and a false positive as equally costly. In regulated settings, they rarely are. Missing a real risk can mean regulatory action, while over-flagging costs a reviewer's time. Vijay's instruction to make the model treat uncertainty as a reason to escalate, and to output a confidence score that triggers human review below threshold, turns the prompt from a vague request into a policy document. That's not prompt engineering. That's risk management expressed in tokens.

The refusal to romanticize the LLM's judgment is what we appreciate most. Vijay isn't saying models are bad. He's saying they're bad at the wrong jobs. The retrieval layer, not the generation step, is where the system's true intelligence lives, because if you retrieve the wrong context, the best model will confidently produce a well-reasoned wrong answer. His point about evaluating retrieval quality separately from final classification accuracy is one we'd shout from the nearest architecture review. Most teams optimize a single F1 score and miss that their retrieval could be great while their generation misweights evidence. And the feedback loop, where human overrides become retrievable context, is the difference between a system that improves and one that repeats the same category of mistake forever. That's the practical takeaway worth quoting: "The more valuable engineering work is deciding what should never touch the model at all."

The open question this leaves us with is cultural, not technical. The cascade approach demands that teams admit the LLM is not the star of the show. It's the escalation path, the expensive specialist you call only when the routine staff, the deterministic rules, and the retrieval index, have done their jobs. That's a harder sell internally than "we built an AI system," but it's the only framing that survives contact with a regulator. If you're building for a high-stakes domain, the first prompt you write shouldn't be for the model. It should be for yourself: which decisions are too important to leave to chance, and which are too routine to waste an inference on? Answer that, and the architecture follows. Get it wrong, and no prompt tweak will save you.

From VentureBeat

Most teams building retrieval augmented generation (RAG) systems for high stakes classification make the same architectural bet: Route every ambiguous case straight to the language model and trust the retrieved context to sort it out. This works fine in a demo. It falls apart the moment the system has to survive an audit, a regulator, or a compliance officer asking why a specific decision was made six months ago.

Read the original at VentureBeat