**Our Take: The Dangerous Distance Between Autonomy and Assurance**
There is a moment in every technology cycle when the speed of adoption outpaces the maturity of the safeguards meant to govern it. We are living in that moment with enterprise AI agents. The data from this wave of Pulse Research delivers a stark, defining image: organizations are handing their agents more independence than they trust the tests to support. Half of all enterprises have already shipped an agent that cleared internal evaluations and then failed in front of a customer. A quarter have seen it happen more than once. That is not a coverage problem, it is a reality-alignment problem. The evaluation said the agent was ready, and it was not. The distance between those two truths is the gap that matters.
What makes this gap so consequential is the direction of travel. Only 5% of enterprises fully trust automated evaluation today, with the most-cited limitation being that evals align poorly with real-world outcomes. Yet two-thirds are either already allowing zero-human-in-the-loop deployment or actively engineering their pipelines to get there within a year. We are building the human out of the decision loop at the exact moment we admit the automated judgment replacing it is not yet reliable. The most common primary tool for evaluation is provider-native evals, tied with having no dedicated tooling at all. A quarter of enterprises run no real-time quality checks on live production traffic. They are watching whether the system is up, but not whether it is right. That is not caution; it is faith dressed up as velocity.
Here is the paradox that should give every technical leader pause: the same organizations removing the human from the deployment decision are simultaneously planning to increase spending on human review workflows. The next dollar is going to observability and people, not just automation. That is not a contradiction, it is an admission. Enterprises sense the gap even as they engineer past it. They are placing a hedge bet on oversight while building toward autonomy, because they know the tests they have cannot yet be trusted to make the calls they are asking them to make. The question is not whether the evaluation layer will consolidate, it will. The market is wide open, and the move toward adoption is already underway. The real question is whether assurance catches up to autonomy before the failures start deploying themselves.
