Amazon AGI director says AI agent reliability, not capability, is blocking enterprise deployment at VB Transform 2026
Our take

The enterprise AI landscape is facing a stark reality check. While enthusiasm for AI agents is undeniably high – Cisco data indicates that a striking 85% of enterprises are currently piloting these tools – the translation to actual production deployment remains stubbornly low, with only 5% making the leap. As Bryan Silverthorn, Director of AGI Autonomy at Amazon, illuminated at VB Transform 2026, the core issue isn’t about chasing ever-more-impressive benchmarks; it’s about reliability. This echoes the findings in "The agent evaluation gap: Enterprise AI organizations have a reality-alignment problem, not a coverage problem — and most are shipping to production anyway," which demonstrates that organizations are granting AI agents more autonomy while trusting the evaluations meant to gate them, suggesting a disconnect between perceived and actual readiness. Furthermore, the challenges of providing AI agents with the necessary context are highlighted in "The AI context gap: Enterprise AI organizations have a trust problem, not a retrieval problem — and most are still building the fix," emphasizing that the infrastructure feeding agents business context is often lagging behind the pace of agent development.
Silverthorn’s framework for understanding AI agent reliability – consistency, robustness, predictability, and safety – offers a compelling lens through which to view this persistent deployment gap. The story of the software QA agent that flawlessly extracted serial numbers for two months before inexplicably failing, due to a subtle visual encoder quirk, powerfully illustrates the point. It’s a reminder that even seemingly robust models can harbor hidden vulnerabilities, and that internal evaluations are often insufficient to replicate the complexities of real-world conditions. This isn't a call to abandon advanced models, but rather a shift in focus. It's about recognizing that model improvement is only *part* of the equation; equally crucial is the development of rigorous measurement processes tailored to the specific risks and stakes of each application. The emphasis needs to move beyond simply checking if an agent *can* perform a task to consistently verifying that it performs it *correctly* over extended periods.
What Silverthorn’s “intern” framework ultimately suggests is a fundamental shift in management approach. Rather than treating AI agents as infallible black boxes, organizations need to manage them with the same care and oversight they would any other employee – albeit one with extraordinary potential and occasional lapses in judgment. This involves proactively identifying potential failure points, building in redundancies and undo capabilities, and consciously defining acceptable risk thresholds. Amazon's willingness to allow agents to occasionally run the wrong experiment in exchange for accelerated research velocity demonstrates a pragmatic approach to balancing innovation with control. This mindset aligns with the evolving security landscape, as discussed in "Zero trust must now move at agent speed," highlighting the need for security architecture to keep pace with the increased autonomy of AI agents.
The future of enterprise AI deployment hinges not on the sophistication of the models themselves, but on the maturity of the organizational processes surrounding them. As Silverthorn rightly concludes, the organizations that ultimately succeed in scaling AI agents won't be those with the most impressive technology, but those with the most adept managers. It’s a compelling reminder that the human element – the ability to anticipate, mitigate, and respond to unexpected outcomes – remains the critical differentiator in the age of AI. The question now becomes: will enterprises embrace this shift in perspective, or will they continue to chase benchmarks while leaving a significant portion of their AI investments languishing in pilot purgatory?
The enterprise AI industry has a math problem. Cisco data shows 85% of enterprises are piloting AI agents, but only 5% have shipped them to production. At VB Transform 2026 on Tuesday, Bryan Silverthorn, Director of AGI Autonomy at Amazon, explained why that gap persists — and why the answer isn't better benchmarks.
Silverthorn, who joined Amazon through its acquisition of Adept AI and now leads multimodal agent training inside the company's AGI lab, argued that reliability must be broken into four distinct dimensions: consistency, robustness, predictability, and safety — a framework he credits to research from Princeton.
"It unpacks different factors that I see tangled together in almost every eval I've ever seen," he said.
Why AI agents pass internal evals but fail real customers in production
The framework matters because agents routinely ace internal evaluations and then collapse in the wild. Silverthorn described a customer that deployed an agent for software QA involving serial number extraction from screens. It worked flawlessly for two months — then began intermittently reading wrong numbers. The culprit: the underlying vision encoder behaved differently depending on where the serial number appeared on screen, and a software change imperceptible to humans triggered the failure.
The lesson, Silverthorn said, is about measurement, not just models. "The models have to be better. Obviously, we're working hard on making the models better," he said. But the deeper takeaway, he added, is that teams need to identify their dimensions of variability and match measurement rigor to the stakes of the application. VentureBeat's own proprietary research, presented before the session, reinforces the point: half of surveyed companies shipped agents that passed internal evals but failed real customers, and enterprises overwhelmingly track uptime while ignoring accuracy — checking the pulse without checking the diagnosis. A related finding underscored how few guardrails exist: most enterprises default to the model makers' own evaluations and little else, leaving their testing strategy, as I described it on stage, a coin flip between trusting the vendor and trusting nothing.
Inside Amazon's 'intern' framework for managing autonomous AI agents
Silverthorn's most memorable prescription was cultural, not technical. Inside Amazon's AGI lab, researchers literally call their agents "interns" — as in, "I'll have my intern talk to your intern." The joke carries a serious operational philosophy. Agents, like interns, are powerful but occasionally clueless, capable of amazing work and spectacular derailment.
Managing them, he argued, requires management skills rather than software skills: asking what could go wrong, adding backups and undo capabilities, and consciously deciding what risk you can accept. "You can ask the intern, 'Hey, what might you do wrong here? How might you mitigate your negative outcomes?'" he said. Amazon's lab has embraced that trade-off, accepting agents occasionally running the wrong experiment in exchange for research velocity — including one agent running experiments around the clock on its own high-level research plan.
What enterprise leaders should do before deploying agents at scale
Silverthorn was candid about the limits of today's technology. Self-improving AI remains "a loaded term," he said — Amazon uses AI to improve its models constantly, but fully autonomous self-improvement is distant. Computer use remains a core focus of his lab, with a commercial trucking customer already using browser automation to stitch together warranty claims across fragmented systems**, though he stressed that no future agent will rely on computer use alone — it will work alongside MCP, APIs, and other tools to complete end-to-end workflows**. And LLM-as-judge techniques, while promising, are just one of several strategies for aligning agent capability with acceptable risk.
For enterprises stuck in pilot purgatory, the path forward starts with a mindset shift: stop asking whether your agent can do something impressive once, and start asking whether it can do it correctly a thousand times in a row.
In other words, the enterprises that escape the 85% ceiling won't be the ones with the smartest agents. They'll be the ones with the best managers.
Read on the original site
Open the publisher's page for the full experience