coding-agent

A fairer benchmark separates agent smarts from system design

Most coding-agent benchmarks collapse the model and its harness into one score, leaving failure causes opaque.

4 min readMachine Learning

The most useful thing about this proposed benchmark is that it refuses to let the model hide behind the harness. Anyone who has watched an agent fail has asked the same question: was the model simply not good enough, or did the surrounding machinery sabotage it? By crossing workflow design with model policy, this setup forces a cleaner answer. It separates the task architecture from the model tier, which is exactly the kind of discipline that has been missing from most coding-agent evaluations. We have seen similar scrutiny applied to vision models in Evaluating AI Vision: A Spider Photo Benchmark Reveals Accuracy Gaps, where raw accuracy numbers obscured where and why models failed. This design takes that same instinct and applies it to agentic systems.

The decision to freeze the original tasks, revisions, tools, retry budgets, and acceptance criteria is the right call. It means the four cells are comparable in a way that most benchmark runs are not. The primary measures also deserve credit for focusing on outcomes rather than process. Cost per independently accepted change and first-pass accepted yield are practical signals. False acceptance and false rejection rates get at the gate quality, which is often where agents quietly succeed or fail. Verification time matters because a benchmark that takes too long to run will not get replicated. These are the numbers that tell a real story. We wrote about the importance of maintaining runnable workflows in Keep Your Data Science Notebooks Running: Six Essential Habits, and the same principle applies here: if the evaluation is not reproducible, the findings are just anecdotes.

The budget normalization problem is the honest weak spot, and the benchmark is right to flag it. Giving every decomposed slice the monolith's full context would subsidize the decomposed condition. That would make the comparison unfair in the other direction. A shared system-level budget is cleaner, but it may obscure which slices actually needed more capacity. That tension is not a flaw in the design; it is a real trade-off that any serious evaluation must face. The fact that the benchmark is wrestling with it before running the experiment, rather than after, suggests this is a benchmark worth watching. It also echoes the analytical rigor seen in Explore the Forrester Function: Beyond Mathematics, a Tool for Machine Learning, where the focus is on understanding the underlying function rather than celebrating surface-level performance.

What we would tell the benchmark is this: do not try to isolate decomposition from model quality too aggressively. Treat the architecture as part of the system. The whole point of an agent is that the model and the harness work together. If you isolate them completely, you are no longer evaluating the thing you would actually deploy. The frontier-decomposed cell is the one to watch. It holds the model tier fixed and changes the task structure. If that cell performs well, it says something specific about how task design can unlock capability. If it performs poorly, it says something equally specific about the limits of decomposition without a model that can exploit it. Either outcome is valuable, and that is the mark of a well-formed experiment. The concrete point to watch is whether the routed decomposed cell can approach the frontier-decomposed cell on first-pass yield. If it does, that is the result worth quoting.

From Machine Learning

I am working on an evaluation design and would appreciate criticism before running it.

Most coding-agent benchmarks collapse the model and its harness into one score. If a run fails, it is difficult to tell whether the cause was model capability, context assembly, task decomposition, tool design, retry policy, or the acceptance gate. A model can also look worse because the harness truncated its output, or look better because the gate only checked for plausible surface markers.

Read the original at Machine Learning