The perfect conversation is a mirage. A single AI agent exchange can score flawlessly on its own merits and still be the clearest signal that a product is fundamentally broken. That tension, laid bare at VB Transform 2026 by leaders from LangChain, Conviva, and CoreWeave, is forcing a long-overdue reckoning in how we measure agentic systems. We have been grading the wrong test. The industry has spent months obsessing over individual trace quality, treating each interaction like a standalone pop quiz, when the real exam is about systemic behavior across an entire user base. As Hui Zhang of Conviva put it, teams are choosing between scalable but ungrounded automated judging and grounded but unscalable human review. That is a false choice, and the path forward demands we stop pretending otherwise.
The shift toward contrastive analysis, comparing cohorts against a baseline rather than scoring traces in isolation, is the most practical insight to emerge from this conversation. Zhang's retail example is the one to remember: a shopper buys a running shoe after a few qualifying questions, and the individual trace looks like a win. But when you zoom out, the clarification ratio for that category is three times higher than baseline, and shoppers are five times more likely to finish their purchase outside the conversation. Those numbers do not lie, and they point to a debuggable, category-specific problem that no single-trace review could ever surface. This is the difference between looking at a single tree and understanding the health of the entire forest. For our readers, this means your evaluation strategy cannot stop at "does this response make sense?" You have to ask "does this response change how users actually behave at scale?" That is a much harder question, but it is the only one that matters. Chase's warning about "eval paralysis" is equally important here; teams that treat evaluation as a one-time gate before launch are building for a world that no longer exists. Evals are the new PRD, a living document that evolves with every launch and every failure. Turlay's experience of reaching 100% test coverage and still shipping bugs is not an anomaly, it is the norm. The answer is not more tests before launch; it is broad, always-on monitoring in production, then using those real-world failures to build targeted offline sets. Start wide, then go deep on what actually breaks.
The move toward cheaper, narrower judge models is a welcome correction, but it needs to be paired with a clear-eyed view of what automation can and cannot do. Turlay's rule, start with the most capable model to prove a task is solvable, then work down, is a pragmatic playbook. Fine-tuning a Qwen model to detect perceived user error and matching Claude Sonnet's performance at a fraction of the cost is exactly the kind of efficiency the market needs. But the human element is not going anywhere. When Turlay talks about signing off on a self-driving car model, and how someone still has to take legal responsibility before deployment, he is touching on a truth that extends far beyond code. The human in the loop is not a bottleneck; it is the source of trust and the mechanism for learning. Chase is right that interactions are essential for the system to improve, and that requires a human presence. The specific detail to watch is how this plays out in regulated industries like finance and healthcare, where the cost of a silent failure is not just a bad metric but a legal liability. The leaders at VB Transform have given us the frame; the burden is on every engineering team to build the contrastive analysis and the always-on monitoring that makes it real. Start comparing cohorts against a baseline today, not because it is trendy, but because your single perfect trace is hiding the very problem that will sink your product tomorrow.
