RAG

Test your RAG apps against unauthorized document access with an open-source tool

A RAG app that quietly retrieves documents a user shouldn't see is a risk most teams don't catch until it's too late.

4 min readMachine Learning

Somewhere in the vast and noisy landscape of AI tooling, a developer has done something quietly useful. They built an open-source checker that tests whether a retrieval-based AI application hands users documents they should never see. The tool runs offline test cases and live HTTP API checks with bearer token or API-key auth. It is small, practical, and aimed squarely at engineers who are tired of guessing whether their RAG pipeline is leaking.

This matters more than the polished demos we usually cover. As we explored in our guide to Unlock LLM Training: A Practical Guide to Distributed Algorithms, the real friction in AI systems is rarely the model itself. It is everything around it: the data plumbing, the access layers, the silent assumptions about who can see what. This tool is a direct response to that friction. It does not promise to fix your architecture or magically harden your app. It simply asks a hard question, repeatedly, and surfaces the answer. That is the kind of honesty we need more of.

Our take? This is the right problem to be solving, but it is also a starting point, not a finish line. Checking whether a retrieval system returns an unauthorized document is a necessary test, but it is not sufficient. Access control in RAG is not just about the retrieval step. It is about what happens after retrieval: how the model synthesizes an answer, whether it can infer hidden facts from multiple harmless documents, and whether your evaluation suite actually covers those cases. The related piece on Exploring Paragraph Structure: How LLMs Navigate Token Space reminds us that the model's internal representations can encode far more than the literal tokens it sees. A tool like this catches the obvious leaks. The subtle ones will still require careful thought.

What we would tell the engineer who asked us directly: yes, try it. Run it against a test environment, not production. Feed it a few cases where you are fairly sure access is broken, and a few where you are not. See if it surprises you. Then, and this is the part that matters, look at what it does not catch. The tool is an invitation to think harder about your own system, not a replacement for doing so. It is also a good reminder that open-source contributions like this move the whole field forward, even when they are small and unpolished. The Unlock ChatGPT for Work: A Practical Guide to Getting Started piece we published earlier makes a similar point: the barrier to useful AI work is rarely the model, it is the discipline of testing and verification around it.

The specific thing we are watching for is adoption. Whether a few engineers run this on a test environment and report back will tell us if the community is serious about access control, or just serious about talking about it. The tool is simple enough to be approachable, which is exactly why it could catch on. But simplicity only gets you so far. The real test is whether it finds something you did not already know. That is the bar. If it does, we will see more tools like it. If it does not, it will quietly join the pile of good intentions. We hope it finds something.

From Machine Learning

I built a small open-source tool that checks whether a RAG application retrieves documents a user shouldn’t have access to.

It supports offline test cases and live HTTP API testing with bearer token/API-key auth.

Read the original at Machine Learning