Enterprises with AI context layers report agent failures at more than twice the rate of those without one
Our take

The recent VB Pulse survey revealing a counterintuitive trend – that enterprises implementing governed AI context layers are *more* likely to report agent failures – underscores a critical, often overlooked, aspect of responsible AI adoption. It’s a stark reminder that simply deploying a technology doesn’t guarantee improved outcomes; it’s the visibility and rigorous assessment of those outcomes that truly matter. The findings, echoing concerns about data governance highlighted in Four of five enterprises that secured AI agent identities still can't contain one that goes rogue, demonstrate that a robust context layer isn’t a magic bullet for eliminating AI errors, but rather a powerful tool for *detecting* them. This shift in perspective – from chasing a flawless record to embracing the process of identifying and correcting inaccuracies – is vital for building trust and ensuring the long-term viability of AI initiatives.
The core issue, as the article rightly points out, isn’t the creation of a context layer itself, but *how* enterprises are approaching context management. The reliance on retrieval over documents, while prevalent, is inherently susceptible to semantic drift and ambiguity, a problem that even sophisticated embedding techniques struggle to fully resolve. The fact that access control and permissions are now driving buying decisions more than retrieval accuracy further highlights this misalignment. The focus is shifting towards simply *getting* data in, rather than ensuring its quality and consistency – a point further amplified by the observation that larger companies, with more data and presumably more scrutiny, report higher failure rates. As noted in SpaceXAI debuts Grok 4.6, overtaking Kimi K3's performance and matching GPT-5.6 Sol for world's third best on Artificial Analysis, the rapid advancement of models is outpacing the development of robust governance frameworks, creating a situation where increasingly powerful AI agents are operating on potentially flawed foundations.
The increase in reported failures with a governed context layer shouldn’t be interpreted as a failure of the layer itself, but rather as a consequence of its effectiveness. These layers provide a crucial reference point for tracing errors back to their source – a broken definition, a stale table, or an inconsistent interpretation. Without such a layer, errors are either masked entirely or attributed to the model’s inherent limitations, hindering the ability to diagnose and rectify the underlying problems. This mirrors the longstanding challenges in data analysis, as Kyle Nesbit aptly describes: the fundamental lack of governed data analysis, now amplified by the scale and speed of AI. The shift in focus, from achieving a pristine failure record to actively monitoring and addressing errors, represents a more mature and realistic approach to AI implementation.
Ultimately, this survey reinforces the need for a more holistic view of AI governance. It’s not enough to simply build a context layer; enterprises must also invest in robust monitoring, validation, and remediation processes. The fact that 79% of enterprises plan to maintain control over their context layers, rather than relying on a single vendor, suggests a growing recognition of the importance of customization and control. As budgets continue to flow into context layer development, the challenge will be ensuring that these investments translate into tangible improvements in data quality and AI accuracy – not just increased visibility of existing problems. The question now becomes: how can enterprises effectively leverage these governed layers to proactively prevent errors, rather than simply react to them?
A company builds a governed context layer specifically to stop its AI agents from confidently giving wrong answers. Once that layer is live, the company is more than twice as likely to report the failure happening — not less.
In the past six months, 68% of enterprises have traced a confident but wrong AI agent answer to missing or inconsistent business context. Thirty-seven percent say it happened more than once, ahead of the 32% who saw it happen only once. The figures come from a VB Pulse July 2026 survey of 101 qualified enterprises with more than 100 employees. That's up from 57% in a VB Pulse survey conducted in June. Recurring failures climbed too, from 31% then to 37% now.
This is the second time VB Pulse has asked enterprises this exact question, once in June and now in July. The failure rate is climbing, not falling, even as more enterprises report a governed layer in production, up from 25% in June to 32% now.
How agents get context determines whether they're wrong
Every AI agent needs some way to know what the business actually means, whether a metric is defined consistently, whether a document is current. That's the operation. The challenge is that enterprises hand agents that context in very different ways, and those ways are not equally reliable.
Retrieval over documents remains the most common approach, the primary source for 31% of enterprises. But a real share of enterprises skip a structured approach altogether. Thirteen percent run agents primarily on long-context loading, feeding documents directly into the model's context window rather than retrieving them. Five percent give agents no structured context at all, just the model's general knowledge. Between them, nearly one in five enterprises are feeding agents business context by brute force or not feeding it at all.
Even the leading approach can still produce a confidently wrong answer. Retrieval works by matching a question to text that looks similar in meaning. Similar wording doesn't guarantee the same meaning. Srijith Rajamohan, an AI research leader at Redis, described exactly this gap in an interview with VentureBeat earlier this year.
"If you have a sentence like 'Rome is closer than Paris' and another that says 'Paris is closer than Rome,' and you do an embedding retrieval followed by a text search, you're not going to be able to tell the difference," Rajamohan said. "The same words exist in both sentences."
Buying shifted to access control. Grading didn't follow.
The way enterprises choose a retrieval system doesn't help close the gap. Access control and permissions now tie ease of data ingestion as the top selection criteria, at 24% each. It's the first time in this survey series that a governance property has led to the buying decision. Retrieval accuracy trails at 15%. The property most directly tied to a confident wrong answer isn't the property most enterprises are buying for.
Once a system is running, correctness is still how enterprises judge it. Response correctness is the primary success metric for 38% of enterprises, twice the next closest answer, security and access control at 19%. Enterprises are shifting how they buy toward governance. They're still grading success on whether the answer is right.
The companies fixing this are the ones reporting it worst
A governed context layer is meant to fix this. It's one shared, agreed-on model of what the business's data means, that every agent and BI tool references instead of guessing on its own. Adoption is far from settled.
Thirty-two percent of enterprises run one in production. Thirty-one percent are piloting or building one right now. Twenty percent are evaluating one. Fourteen percent have no plans to, and 4% don't know.
Compare that adoption data against who's actually had the failure, and the picture inverts. Among the 91 enterprises able to say whether they'd experienced the failure at all, those running or building a governed layer report it recurring at 50%. Those without one report it at 21%.
A governed layer doesn't cause the failure — it's what makes the failure visible in the first place. Tracing a bad answer to a broken definition or a stale table requires a shared, governed reference point. A context layer provides that. Without one, the same wrong answer still happens — it just gets chalked up to the model, or never gets traced at all.
The pain point predates AI by decades. Kyle Nesbit, founder of the semantic layer startup Credible Data, described it to VentureBeat last month. "It's the same pain point people have had for 30 years, the lack of governed data analysis," Nesbit said. "Now with AI, it's the same problem, but orders of magnitude more chaos and pain."
Company size sharpens the same point. Enterprises with more than 1,000 employees report recurring failures at 55%, against 30% for those between 101 and 1,000 employees. That's despite the bigger companies being less likely to have a layer already in production, 24% against 37%. More instrumentation and more people asking why a number was wrong turns up more failures, not fewer. A clean record is not evidence of a healthy context layer. It's at least as likely to be evidence that nobody's checking.
What this means for enterprises
Here's what this adds up to for enterprises building on this layer.
Retrieval alone will not close the context gap. RAG remains the default context source, and nearly one in five enterprises are running agents on long-context loading or no structured context layer at all. More documents or a bigger index doesn't fix a definition that means two different things in two different systems.
The budget is moving faster than the infrastructure is shipping. Sixty-three percent of enterprises are already building or running a governed context layer. Only 32% have actually gotten one into production. That gap is where the spend is going, not where the problem has been solved.
A clean failure record is a red flag, not a green one. The 22% of enterprises reporting no context failure at all are not the best-governed group. They're the group least likely to be checking. The size data backs this up directly. Larger enterprises report recurring failures at nearly twice the rate of mid-market peers, despite being less likely to have a governed layer in production, not more.
No one is planning to hand the layer to a single provider. Seventy-nine percent of enterprises intend to keep at least part of the context layer outside any one vendor's stack, split between best-of-breed tools and an explicit mix. Just 12% plan to consolidate onto a single provider's native context stack.
The finding that organizations aren't likely to hand over control to a single provider is a theme that VentureBeat has reported on consistently this year. Michael Ni, an analyst at Constellation Research, put it bluntly earlier this year when DataHub's context layer push first landed.
"Whoever controls runtime context, controls the AI decision layer for enterprise data," Ni said.
Read on the original site
Open the publisher's page for the full experience