automated anomaly detection

Why enterprises double down on AI automation after being burned

Enterprises that have watched an AI agent pass its evals and then fail in front of customers aren't retreating from automation, they're moving faster toward removing humans from deployment decisions.

4 min readVentureBeat
Why enterprises double down on AI automation after being burned

**Our Take: The Automation Paradox at the Heart of Enterprise AI**

The latest VB Pulse data reveals a contradiction that should worry every data leader, but it isn't the one you might expect. We keep hearing about the evaluation gap, the distance between what an AI agent promises in testing and what it delivers in production. That gap remains stubbornly wide, with 49% of surveyed enterprises reporting that an AI feature cleared internal checks before disappointing a customer. But the more troubling signal is how companies respond to that failure. Instead of pulling back on autonomy, 85% of those burned organizations are accelerating toward removing humans from deployment decisions entirely. They are not slowing down because they lack trust in the technology. They are speeding up because they have shifted their trust to a different layer of the stack.

This is the paradox of confidence in automated evaluation. In July, 13% of enterprises said they trust automated checks, up from just 5% in June. Yet the very organizations that experienced a test-passing agent fail in front of customers are six times less likely to place complete faith in those checks. That split makes intuitive sense. But here is what should give you pause: those same burned enterprises are also the most aggressive in removing human approval from release pipelines. They are not doing this out of recklessness. They are doing it because they believe the bottleneck is no longer the pre-deployment test, it is the speed of iteration. As Raindrop.ai CTO Ben Hylak notes, the Fortune 100 are deprioritizing eval maintenance entirely, leaning instead on anomaly detection before and after production. The shift is from preventing every failure to catching it faster downstream. That is a fundamentally different operating model, and it carries real risk.

The data on production monitoring exposes the soft underbelly of this approach. While 67% of enterprises are moving toward no-approval deployment, only 28% of that group automatically checks whether live outputs are semantically correct. Most are still watching gateway metrics like latency and cost, signals that catch outages but miss a fluent, confident, and wrong answer. The AI agents shared user images incident earlier this year was a stark reminder that a system can pass every internal check and still fail spectacularly in the wild. What the VB Pulse survey shows is that enterprises are preparing for that reality by building human review as a backstop, not a gate. The burned group is increasing investment in people-centered review faster than any other category. They are not removing humans from the loop entirely. They are moving them from the front door to the emergency exit.

The practical takeaway here is uncomfortable but clear: **if you have experienced an AI failure in production, you are more likely to remove human approval from deployment, not less.** That is not a bug in the data. It is a reflection of maturity. The organizations that have seen evals fail understand that a passing score is a starting point, not a certification of reliability. They are spending on integration and observability because they know the next failure is a matter of when, not if. The question that remains is whether production monitoring can catch up before the volume of autonomous deployments outpaces the human backstop. Watch the budget lines: if people-centered review spending continues to climb alongside no-approval deployment, we are building a system that automates speed and relies on humans to clean up the mess. That may work at 100 agents. It will break at 10,000.

From VentureBeat

Enterprises that already got burned by an AI agent passing its evals and then failing in production are moving faster toward removing humans from deployment decisions, not slower — even as trust in automated evaluation is rising across the board, new VB Pulse research shows.

In July, 13% of 108 enterprises surveyed said they trust automated evaluation, up from just 5% the month prior. Meanwhile, survey respondents citing poor alignment between tests and real-world results as their biggest concern fell 10 points, from 29% to 19%, month over month.

Read the original at VentureBeat