Multi-Document RAG

Unlock hidden connections across unrelated documents with intelligent outlines

A folder of unrelated PDFs doesn't need a shared index to be searchable.

3 min readTowards Data Science
Unlock hidden connections across unrelated documents with intelligent outlines

Most approaches to multi-document retrieval start with a shared schema: if the files don't share fields, the reasoning goes, you can't build an index. That assumption is flipped. Instead of forcing unrelated PDFs into a uniform structure, it treats the folder as a single long document with a nested outline, where each file contributes one summary line and its own table of contents. Retrieval then routes down two levels: first to the file, then to the specific section. That is a quietly radical idea, because it removes the need for a common key while still giving the retriever a clear path to follow.

For anyone who has wrestled with enterprise document dumps, this is the practical unlock. Most folders are not curated datasets; they are a pile of invoices, manuals, and reports that share only a topic and a deadline. The related work in Exploring Paragraph Structure: How LLMs Navigate Token Space shows that structure matters at the token level, and the same logic extends to the document level. A nested outline is not a workaround; it is a way to give the model a map without forcing every page into the same template. The summary line per file is the anchor, and the table of contents does the heavy lifting for navigation. That is a design choice worth stealing.

What we would tell a reader who asks, "Should I use this?" is: only if you stop thinking about retrieval as a search problem and start thinking about it as a routing problem. The method does not try to answer a query by scanning every page. It narrows the space first, then drills down. That is closer to how a human handles a messy folder, and it is why the approach feels robust. It also complements the work in Bridging Retrieval and Action: A New Approach to AI Tasks, where retrieval is only half the pipeline. Here, the retrieval step is deliberately simple, but it is designed to feed a downstream task cleanly. The two articles agree on a core point: complexity should live in the structure, not in the query logic.

The open question is how well the nested outline scales when the folder grows into the hundreds of files. A two-level route assumes the first summary line is accurate enough to pick the right file, and that each table of contents is clean. In the real world, PDFs are messy, and a single misleading summary can send retrieval down the wrong branch. That is the tradeoff. But the more interesting takeaway is that the approach avoids the usual plea for better metadata. It works with what exists: a filename, a summary, a table of contents. That is a practical standard most teams can meet. If you are evaluating RAG approaches, the specific question to ask is not whether your documents share fields, but whether you can write one true sentence about each file. If you can, the nested outline gives you a retrieval path worth testing.

From Towards Data Science

Enterprise Document Intelligence [Vol.1 #14B] - No shared fields means no index to build. One summary line per file plus each file’s own table of contents, and retrieval routes down two levels

The post Multi-Document RAG: A Folder of Unrelated PDFs Is One Long Document with a Nested Outline appeared first on Towards Data Science.

Read the original at Towards Data Science