enterprise data management

A single filter closes the retrieval gap in Azure OpenAI email automation.

Closing an Azure OpenAI assistant's retrieval gap didn't require a new identity platform.

4 min readVentureBeat
A single filter closes the retrieval gap in Azure OpenAI email automation.

The evaluation scores were clean, and the unit tests passed. None of them asked the question that mattered. Egiziago Cioffi built an Azure OpenAI email assistant that auto-resolves 60% of inbound customer email, and his team's entire quality process confirmed what they wanted to believe: the agent worked. Then he ran a low-privilege account against the same questions a high-privilege account had already asked, and the outputs diverged. The assistant returned SharePoint content the low-privilege user could not have opened directly. That is not a subtle bug. That is a structural collapse of authorization boundaries, hidden in plain sight because the evaluation framework never thought to look. We have said before that retrieval is not evidence, and here is the proof: the assistant answered correctly, and it was still wrong. The answer was accurate. The permission boundary was not. Cioffi's fix did not require a new identity platform, and that is the part worth pausing on. The industry is spending tens of billions of dollars on identity governance acquisitions like CrowdStrike's $740 million purchase of SGNL and Palo Alto Networks' $25 billion CyberArk deal, both closing in February 2026. Those platforms govern which service accounts exist, what they can reach, and when their tokens expire. They do not govern the retrieval permission boundary. That is the moment a correctly scoped service account fetches content on behalf of a user who holds fewer permissions than the indexing job does. Every credential in that chain is legitimate. The service account is clean. The knowledge base is correctly indexed. A low-privilege user queries the assistant, and it answers from the full indexed scope. Nothing flags the retrieval because no credential was misused. Cioffi closed the gap with a query-time filter that checks the requesting user's SharePoint permissions before the model sees a chunk. That filter narrowed what the assistant could reach, and it still auto-resolves roughly 60% of inbound email. The tradeoff is real and describable: some content the assistant previously used is now excluded because the requesting user's permissions do not reach it. The deeper problem is that answer-quality evaluations are the wrong tool for this job. Cioffi's team tested factual accuracy, relevance, and task completion. They did not test whose permissions the retrieval pipeline uses when it fetches source material. That question is not in the evaluation framework. Straiker's red team ran more than 1,700 successful exploit attempts against production agents and found that 91% of successful attacks ended in silent data exfiltration with no malware required. The U.K. AI Security Institute documented 19 unsanctioned agent actions in a permissive test environment. Neither of those findings is specifically about retrieval entitlements, but they share the same root cause: agents act outside the scope their deployers intended, and no runtime check catches the deviation before it causes damage. The fix is not another security product. The fix is a two-account test that takes thirty minutes. Run the same question a high-privilege account has already put to the assistant, then compare the output against what the low-privilege account can access directly. If the assistant returns more than the account's direct access would allow, the retrieval permission boundary is not enforced at query time. That test costs two accounts and half an hour. It produces a result an evaluation score cannot replicate. The question to watch is not whether Cioffi's filter is the right pattern, because it is a solid pattern. The question is how many production deployments are running on custom Azure OpenAI pipelines that bypass the native ACL trimming layer entirely. Azure AI Search ships document-level ACL trimming via Entra-based tokens, and SharePoint ACL sync followed in preview. But that capability does not cover every deployment path, and Microsoft's own documentation states that if the permitted-groups field is not mapped, document-level access is disabled. That is a fail-open default in a first-party path. Cioffi's deployment took the custom-pipeline route, and his logs are the evidence that the gap survives every evaluation his team ran. The takeaway we would give any reader is direct: run the two-account comparison before the next deployment goes live.

From VentureBeat

Egiziago Cioffi is the IT and Enterprise Architect and CEO of SynSphere Italia, a Microsoft partner based in Milan. He built an agent himself. He wrote the indexing job, configured the Azure OpenAI retrieval pipeline, connected it to SharePoint, and watched it pass every evaluation his team ran.

His Azure OpenAI email assistant auto-resolves about 60% of inbound customer email, Cioffi told VentureBeat in written responses to our interview questions. The evaluation scores were clean, and the unit tests passed. None of them asked the question that mattered.

Read the original at VentureBeat