The premise that a case file's index should be defined before a single PDF is opened is the most practical reframing of retrieval-augmented generation we have seen in a while. Most teams treat RAG as a fire-and-forget search problem: chunk the documents, stuff a vector store, and hope the model stumbles onto the right answer. The argument is sharper than that. It says the relational tables, the schema of what matters for a given case type, are the actual scaffolding. The PDFs are just the raw material. This aligns with something we have been watching across our own coverage, like how Exploring Paragraph Structure: How LLMs Navigate Token Space shows that token position is a coordinate system, not a container. If you do not define the coordinates first, you are just guessing at the map.
The two questions worth building for are not retrieval questions at all. They are analytical ones: what does this case require, and what is missing? That is a subtle but decisive shift. Retrieval assumes the answer exists in the folder. The better question assumes the folder is incomplete, and the index is what tells you what is absent. This turns RAG from a lookup tool into a diagnostic one. For any team building on top of LLMs, this is the difference between a system that regurgitates and one that reasons. We would tell a reader who asked us about this: stop optimizing for recall on the documents you have, and start designing for the relations that should exist across them. The index is the product. The PDFs are just the payload.
This also connects to a broader tension we have explored in Bridging Retrieval and Action: A New Approach to AI Tasks. That work separated retrieval from action, then wired them together explicitly. This does something similar at the data layer: it separates the schema from the content. The schema is not derived from the content; it precedes it. That is a meaningful design choice. It means the system knows what it is looking for before it looks, which is exactly how a human paralegal actually works. They do not read every page to find out what matters. They know the case type, the required documents, the likely gaps, and they go looking for those specifically. The AI version of that is not a bigger context window. It is a pre-defined relational table.
The practical takeaway, and the one we would quote, is this: the index is the retrieval strategy, not the search query. If you build the relational table for the case type, you are not asking the model to find answers; you are asking it to fill in a known structure. That is a profoundly more reliable task. The open question we are left with is whether teams will make the investment to define those schemas, or whether they will keep chasing better embeddings on messy folders. The next wave of enterprise AI will not be won by fancier recall. It will be won by teams that decide what the case demands before the folder is even opened. That is the discipline worth building for, and it is the one most projects still skip.
