rows.com

Why enterprise AI agents stall at deployment, not capability

At VB Transform 2026, Bryan Silverthorn from Amazon made a pointed case: enterprise AI agents aren't failing because they're not smart enough.

4 min readVentureBeat
Why enterprise AI agents stall at deployment, not capability

The 85% to 5% collapse between pilot and production is the kind of stat that should worry anyone building on top of large language models, but Bryan Silverthorn's diagnosis at VB Transform 2026 reframes the problem in a way that actually helps. He is not asking for better benchmarks or bigger models, at least not first. He is asking teams to admit that the gap is about reliability, and that reliability is not a single switch you flip. It is four distinct dimensions, consistency, robustness, predictability, and safety, that usually get tangled together in every internal eval. That framing alone is worth taking back to your engineering standup this week.

The story about the serial number extraction agent is the one to hold onto. It worked flawlessly for two months, then started reading wrong numbers because a vision encoder shifted behavior based on where the number appeared on screen. A software change that no human would notice triggered the failure. This is the difference between a demo and a deployment, and it aligns with what Morgan Stanley and others have been learning about modernizing complex systems: you cannot rely on static checks when the environment shifts under you, whether you are orchestrating AI agents with open-source tools or scaling AI workflows with architecture as code. The same discipline applies: you need to identify your dimensions of variability and match measurement rigor to the stakes. Half of the companies VentureBeat surveyed shipped agents that passed internal evals and then failed real customers. That is not a model problem. That is a management problem.

Silverthorn's "intern" framing is the most practical thing said on that stage. Agents are powerful, occasionally clueless, and capable of spectacular derailment, exactly like a good intern. The operational philosophy that follows is not cute; it is actionable. You ask the agent what it might do wrong. You add backups and undo capabilities. You decide consciously what risk you can accept. Amazon's lab accepts that its agents occasionally run the wrong experiment in exchange for research velocity, including one agent running experiments around the clock on its own high-level plan. That is a deliberate trade-off, not an accident. Most enterprises are not making that trade-off consciously. They are defaulting to the model maker's own evaluations and little else, which means their testing strategy is a coin flip between trusting the vendor and trusting nothing. If you are stuck in pilot purgatory, the question is not whether your agent can do something impressive once. The question is whether it can do it correctly a thousand times in a row.

The takeaway to quote: "Stop asking whether your agent can do something impressive once, and start asking whether it can do it correctly a thousand times in a row." That is the mindset shift that separates the 5% from the 85%. The enterprises that escape the ceiling will not be the ones with the smartest agents. They will be the ones with the best managers. And if you are leading one of those teams, the immediate next step is to audit your own evaluation practices against Silverthorn's four dimensions, specifically, ask which of the four you are ignoring. Because the math says you are probably tracking uptime while ignoring accuracy, and that is checking the pulse without checking the diagnosis. Watch for Amazon's own agent work in computer use and browser automation to expose just how many enterprise workflows can be stitched together with enough management discipline, and how quickly that changes the competitive bar for everyone else.

From VentureBeat

The enterprise AI industry has a math problem. Cisco data shows 85% of enterprises are piloting AI agents, but only 5% have shipped them to production. At VB Transform 2026 on Tuesday, Bryan Silverthorn, Director of AGI Autonomy at Amazon, explained why that gap persists — and why the answer isn't better benchmarks.

Silverthorn, who joined Amazon through its acquisition of Adept AI and now leads multimodal agent training inside the company's AGI lab, argued that reliability must be broken into four distinct dimensions: consistency, robustness, predictability, and safety — a framework he credits to research from Princeton.

Read the original at VentureBeat