4 min readfrom VentureBeat

Amazon AGI director says AI agent reliability, not capability, is blocking enterprise deployment at VB Transform 2026

Our take

Amazon’s Bryan Silverthorn, Director of AGI Autonomy, recently pinpointed a critical obstacle hindering enterprise AI agent deployment: reliability, not inherent capability. Addressing attendees at VB Transform 2026, Silverthorn highlighted a concerning trend – 85% of enterprises pilot AI agents, yet only 5% reach production. His framework, emphasizing consistency, robustness, predictability, and safety, underscores the need for rigorous measurement, echoing findings that many agents fail after initial evaluations.
Amazon AGI director says AI agent reliability, not capability, is blocking enterprise deployment at VB Transform 2026

The enterprise AI landscape is currently grappling with a significant disconnect: a surge in AI agent piloting alongside a startlingly low rate of production deployment. Cisco data paints a clear picture – 85% of enterprises are experimenting with AI agents, yet only 5% are confident enough to release them into live environments. As Bryan Silverthorn, Director of AGI Autonomy at Amazon, articulated at VB Transform 2026, the issue isn't simply about building more capable models; it's about ensuring reliability. This echoes recent findings, as highlighted in [Agentic orchestration: Enterprise AI organizations have a deployment problem, not a platform problem — and most are calling chatbots agents], where a platform consolidation is occurring, indicating a shift toward manageable deployments. Furthermore, Stripe's recent benchmark, detailed in [Stripe Benchmark Shows AI Agents Build Integrations but Struggle with Validation], reveals that while agents excel at integration, validation remains a considerable hurdle, reinforcing the need for robust reliability measures. Thinking Machines' open-sourcing of Inkling, as discussed in [Thinking Machines open sources first multimodal language model, Inkling, focused on low cost and 'resistance to censorship’], presents an interesting avenue for enterprises seeking greater control over their agent deployments, but the underlying reliability challenges persist regardless of model architecture.

Silverthorn’s breakdown of reliability into consistency, robustness, predictability, and safety offers a crucial framework for understanding this problem. The anecdote about the software QA agent that initially performed flawlessly before intermittently misreading serial numbers is a powerful illustration of how easily these agents can fail in real-world conditions. This failure wasn't due to a fundamental flaw in the model itself, but rather a subtle variation in the input data – a change in the serial number's position on the screen that triggered unexpected behavior in the vision encoder. This underscores the point that measurement and rigorous testing are paramount, and that organizations must move beyond superficial evaluations to identify and mitigate sources of variability. The emphasis on matching measurement rigor to the stakes of the application is particularly astute; a minor inaccuracy in a customer service chatbot is far less consequential than an error in a financial trading algorithm.

The most compelling aspect of Silverthorn’s presentation was his cultural prescription: treating AI agents like "interns." This seemingly playful analogy highlights the need for management oversight and risk mitigation. Just as a manager wouldn't blindly trust an intern to execute a critical task without supervision, organizations shouldn't deploy AI agents without robust backups, undo capabilities, and a clear understanding of potential failure modes. Embracing this mindset shift – from focusing solely on capability to prioritizing reliability – is arguably the biggest barrier to widespread enterprise adoption. It’s a recognition that even the most advanced AI agents are, at their core, tools that require careful management and human oversight. Amazon's own acceptance of occasional experimental errors in exchange for research velocity demonstrates a pragmatic approach to balancing innovation with responsible deployment.

Ultimately, the future of enterprise AI hinges not on the development of ever-more-powerful models, but on the ability to reliably manage and integrate those models into existing workflows. The 85% ceiling represents a significant opportunity for organizations that can master this challenge. As AI agents become increasingly prevalent, the demand for skilled “AI managers” – individuals who can assess risks, implement guardrails, and ensure consistent performance – will only continue to grow. The key question moving forward is: will enterprises invest in developing these critical skills and establishing robust management frameworks, or will they remain trapped in a perpetual state of pilot purgatory?

The enterprise AI industry has a math problem. Cisco data shows 85% of enterprises are piloting AI agents, but only 5% have shipped them to production. At VB Transform 2026 on Tuesday, Bryan Silverthorn, Director of AGI Autonomy at Amazon, explained why that gap persists — and why the answer isn't better benchmarks.

Silverthorn, who joined Amazon through its acquisition of Adept AI and now leads multimodal agent training inside the company's AGI lab, argued that reliability must be broken into four distinct dimensions: consistency, robustness, predictability, and safety — a framework he credits to research from Princeton.

"It unpacks different factors that I see tangled together in almost every eval I've ever seen," he said.

Why AI agents pass internal evals but fail real customers in production

The framework matters because agents routinely ace internal evaluations and then collapse in the wild. Silverthorn described a customer that deployed an agent for software QA involving serial number extraction from screens. It worked flawlessly for two months — then began intermittently reading wrong numbers. The culprit: the underlying vision encoder behaved differently depending on where the serial number appeared on screen, and a software change imperceptible to humans triggered the failure.

The lesson, Silverthorn said, is about measurement, not just models. "The models have to be better. Obviously, we're working hard on making the models better," he said. But the deeper takeaway, he added, is that teams need to identify their dimensions of variability and match measurement rigor to the stakes of the application. VentureBeat's own proprietary research, presented before the session, reinforces the point: half of surveyed companies shipped agents that passed internal evals but failed real customers, and enterprises overwhelmingly track uptime while ignoring accuracy — checking the pulse without checking the diagnosis. A related finding underscored how few guardrails exist: most enterprises default to the model makers' own evaluations and little else, leaving their testing strategy, as I described it on stage, a coin flip between trusting the vendor and trusting nothing.

Inside Amazon's 'intern' framework for managing autonomous AI agents

Silverthorn's most memorable prescription was cultural, not technical. Inside Amazon's AGI lab, researchers literally call their agents "interns" — as in, "I'll have my intern talk to your intern." The joke carries a serious operational philosophy. Agents, like interns, are powerful but occasionally clueless, capable of amazing work and spectacular derailment.

Managing them, he argued, requires management skills rather than software skills: asking what could go wrong, adding backups and undo capabilities, and consciously deciding what risk you can accept. "You can ask the intern, 'Hey, what might you do wrong here? How might you mitigate your negative outcomes?'" he said. Amazon's lab has embraced that trade-off, accepting agents occasionally running the wrong experiment in exchange for research velocity — including one agent running experiments around the clock on its own high-level research plan.

What enterprise leaders should do before deploying agents at scale

Silverthorn was candid about the limits of today's technology. Self-improving AI remains "a loaded term," he said — Amazon uses AI to improve its models constantly, but fully autonomous self-improvement is distant. Computer use remains a core focus of his lab, with a commercial trucking customer already using browser automation to stitch together warranty claims across fragmented systems**, though he stressed that no future agent will rely on computer use alone — it will work alongside MCP, APIs, and other tools to complete end-to-end workflows**. And LLM-as-judge techniques, while promising, are just one of several strategies for aligning agent capability with acceptable risk.

For enterprises stuck in pilot purgatory, the path forward starts with a mindset shift: stop asking whether your agent can do something impressive once, and start asking whether it can do it correctly a thousand times in a row.

In other words, the enterprises that escape the 85% ceiling won't be the ones with the smartest agents. They'll be the ones with the best managers.

Read on the original site

Open the publisher's page for the full experience

View original article