The CFO's decision wasn't a failure of the AI. It was a failure of imagination, specifically the imagination that built the eval harness in the first place. An agent aced every metric we tend to celebrate, then got canned because its successful resolutions cost more than the humans it replaced. That's the story we need to sit with, because it reveals a hard truth: our evaluation culture is still obsessed with task completion while the people paying the bills are looking at unit economics.
We've all been guilty of building systems that optimize for the wrong thing. The agent wasn't broken. It did what it was asked to do, which was to solve problems. But no one asked the question that actually mattered in production: did the cost of each resolution, including the time spent handling edge cases, the integrations, and the rework, beat the cost of a human doing the same job? The answer was no. This is where we see the gap between Explore how AI agents learn by editing context, not model weights and the reality of deployment. The former is fascinating intellectually, but the latter demands we think about how agents interact with messy, real-world workflows where the marginal cost of a mistake isn't a point on a graph, it's a budget line item.
Our take is blunt: if you are building AI agents, your eval harness is lying to you if it doesn't include a cost-per-outcome metric. We're not talking about accuracy or resolution rate. We're talking about the total cost of ownership. The author learned this the hard way, and we should all be grateful for that lesson. It's not enough to measure if the agent can do the task; you must measure whether it's worth doing. This is the same logic that applies to Evolve Your Recommendations: Real-World Insights on Adaptive Systems where we see that the real complexity isn't the model, but the feedback loops around it. The agent's success in a sandbox is irrelevant if the feedback loop in production is a black hole of cost.
So what do we tell a reader who comes to us with a similar story? Stop optimizing for the eval. Start optimizing for the business case. The specific takeaway here is that your next eval should include a line item for the cost of a human review for every unresolved case, and a cap on the average cost per successful resolution. If your agent can't beat that number, it doesn't matter if it passes every other test you can throw at it. The CFO will always have the final say, and they're not grading on a curve. They're grading on the P&L. The open question we're left with is whether we can build a new generation of eval harnesses that measure economic viability with the same rigor we apply to accuracy, or if we'll keep watching technically perfect agents get killed by the one metric that truly matters.
