The shift Amatriain is describing is not a tweak to how we build software. It is a fundamental reordering of the product development lifecycle, one that puts the test before the code, the outcome before the feature. When he says that evals are the new PRD, he is telling us that the hours we used to spend arguing about requirements are better spent defining what success looks like in measurable, adversarial terms. For teams still treating evaluation as a final checkpoint before launch, this is a direct challenge to their entire workflow. The Talking to My AI Clone Taught Me to Question the Tech piece we ran recently reminds us that our own reactions to AI outputs are unreliable; if we cannot trust our gut on a single interaction, how can we trust it to define a system's correctness at scale? Amatriain's answer is that we stop trusting the gut entirely and start trusting the eval suite.
The data from our own VB Pulse research makes the urgency concrete, and it is not comfortable. Sixty-six percent of enterprises are already shipping production AI without consistent human review, yet only five percent trust the automated evals that are supposed to catch failures. That gap is not a lag in adoption; it is a structural flaw in how evaluation is being implemented. Half of the teams we surveyed have watched an agent sail through internal tests and then fail in front of a real customer. That is not a bug report, that is a product definition. If you are building an AI feature today, your PRD is not the spec document your PM is drafting, it is the suite of edge cases you have encoded into your evaluation harness. The teams that internalize this will ship with confidence; the ones that treat it as a buzzword will keep paying for the lesson in production incidents and user churn. This is why our Verify Your AI's Understanding: A Simple Check for Tax Season piece resonated so strongly, because it shows that even a basic eval, one that checks for factual grounding, is more valuable than any amount of prompt engineering.
But there is a tension in Amatriain's argument that we should not gloss over. He frames guardrails as a necessary evil, a bias that corrupts the feedback loop, and he wants to minimize their footprint over time. That is a bold position, and it is also a risky one. The counterargument, raised by other speakers at Transform, is that for high-risk actions like booking travel or handling financial transactions, the feedback loop itself is the problem. If a user feels pressured to correct an agent that has already made a mistake, the feedback is not clean signal, it is contaminated data. We would tell a reader to take Amatriain's toll gate approach seriously, but to be honest about the fact that his three-layer governance model, principles, processes, automation, only works if the principles are actually communicated and the processes have teeth. The fact that 54 percent of enterprises have already had an agent security incident should be a cold shower for anyone who thinks this is a theoretical exercise. The Navigating AI/ML Job Requirements: A Shift in Expected Skills article we published highlights a related problem: the market is already demanding people who can build and evaluate these systems, but the skill set is still poorly defined. The next attacker is not a script kiddie, it is another agent designed to probe your system's blind spots, and the time to fix those blind spots is before you write the code, not after the incident report lands.
