ThinkingBox

Why one successful agent run doesn't mean the database agrees

A single successful agent run can hide a uncomfortable truth: the database may never have agreed with what the agent claimed to do.

4 min readMachine Learning
Why one successful agent run doesn't mean the database agrees
ThinkingBox: Solving an agent task once vs. solving it 20/20: 507 stateful workflows graded on terminal database state [R]

A single successful agent run tells you very little about whether the agent can do the job twice, and Thinkingbox-Bench just proved it with numbers we should all sit with. Microsoft's new benchmark ran 507 enterprise workflow tasks across 20 independent attempts per model, grading not on whether the agent *said* it finished but on whether the backend database actually ended up in the correct state. The result is a near-reversal of the leaderboards we're used to seeing: Kimi-K3 solved 93.89% of tasks at least once, but only 13.41% of tasks on all 20 attempts. Claude Opus 5 discovered fewer tasks (79.09%) yet repeated success on 47.53% of them. Pick your metric and you pick your winner. That is not a footnote; that is the story.

This matters because most of us are still benchmarking with a single pass and calling it a day. The paper's retrospective ablation over 121,680 trials found that 67.24% of failed attempts still terminated cleanly, invoked a state-changing tool, and ended without an error. A completion-style proxy would have scored those as done. The state checks then revealed the truth: 77.61% had wrong field values, 43.30% had unintended extra effects, and 25.36% were missing required effects. If your evaluation stops at "the agent stopped without crashing," you are not measuring reliability; you are measuring politeness. This connects directly to the trap we flagged when Gemini Argon Tops the Charts but Skips the Test That Counts, where headline scores obscure what actually matters for deployment. And it echoes the caution from How to benchmark online AI models without losing your data to training: the way you set up the test determines what you learn, and a leaky or shallow setup teaches you the wrong lesson.

The deeper point is that discovery and repeatability are different capabilities, and conflating them is how you end up trusting a model that works once. For anyone building on top of these systems, the practical question is not "can it do the task?" but "can it do the task every time, from a clean state, without leaving the database in a mess?" Thinkingbox's all-20 metric is an observed count, not a guarantee, but it is a far more honest signal than pass@20 for production decisions. The authors ask whether a reliability leaderboard should emphasize the fixed-campaign count or a plug-in estimate. We would argue both, side by side, because they answer different questions: one records what actually happened, the other projects what might happen under assumptions. Hiding either invites the exact overconfidence this benchmark exposes.

The concrete thing to watch is how the field responds. If model providers start reporting all-20 alongside pass@1, we will finally have a language for discussing reliability instead of raw capability. If they resist, ask why. Meanwhile, the benchmark is open and runnable through the Hugging Face OpenEnv environment, which means you can test your own model before you trust it. That is the fastest way to disagree with the authors using your own numbers, and it is the only way to know whether your agent's success is a habit or a fluke. The data is public. The question is whether you will look before you deploy.

From Machine Learning

Disclosure: I'm one of the authors (Microsoft). The paper, code, dataset are public and ThinkingBox is on Hugging Face OpenEnv as well. Raw evaluation trajectories are not released Links at the bottom.

We wanted to know how much of a single agent success rate survives repetition, and whether "the agent finished the task" means the backend/database actually ended up in the correct state.

Read the original at Machine Learning