The gap between Alibaba's marketing table and VulcanBench's independent run is not a scandal, but it is a warning. Both results are real, and both are defensible. Alibaba gave Qwen 3.8-Max a five-hour timeout on coding tasks and up to twelve hours on PaperBench; VulcanBench allowed at most sixty minutes of wall clock. When a model spends most of its token budget on reasoning, the difference between a solved task and a silent failure is often just a clock. That is why the price-per-token comparison everyone published last week tells you less than it used to. Unlock LLM Training: A Practical Guide to Distributed Algorithms draws a parallel in distributed systems: you cannot optimize what you cannot measure, and here the missing measurement is time.
We have spent years treating benchmark scores as if they were intrinsic properties of a model, when they are really properties of a budget. The honest fix is to stop asking which model is smarter and start asking what a successful task costs you. That means computing cost per successful task: total spend, including every failed attempt and every empty response, divided by the tasks that actually passed your acceptance check. VulcanBench already reports dollars per solved task as a headline column. Long-Horizon-Terminal-Bench publishes per-task cost next to accuracy, and its most instructive row shows GPT-5.4 at roughly $26 per task with a much lower pass rate than Grok 4.5 at about $11. Vendors are already moving this way, HubSpot charges 50 cents per resolved conversation, Zendesk bills per automated resolution, and Fin charges 99 cents per outcome. The market is converging on this number because it is the only one that predicts your invoice.
The deeper problem is that your failure rate is partly a configuration setting, and almost no one is tracking it. Long-Horizon-Terminal-Bench ran 17 frontier models through a shared harness with one 90-minute attempt per task. Timeouts accounted for 79% of unresolved runs, against 19% for agents that stopped on their own and 3% for harness errors. That is the clearest evidence yet that benchmarks are implicitly measuring time efficiency, whether they admit it or not. The same mechanism appears in Claude Opus 5's own results: its lowest-effort setting solved 20 of 23 tasks, while high effort solved 18. High effort produced the fewest wrong answers, but it ran out of clock instead, and a timeout scores zero. For a routing ladder, the standard assumption is that escalating to more reasoning when a cheap attempt fails is a safe bet. For a meaningful share of model and task combinations, that assumption is wrong, and you pay the higher rung's price to escalate into a timeout.
So here is what we would tell a reader who asks what to do this week. Emit a failure reason on every agent run as a required field, with budget exhaustion, verifier failure, and harness error as distinct values. Until you can separate a timeout from a wrong answer, your pass rate is measuring two things at once and you cannot tell which one to fix. Compute cost per successful task per effort level, not just per model, and cap on tokens rather than wall clock unless latency is genuinely in your service level objective. A wall-clock cap scores your provider's serving speed as model quality. Finally, check the default effort setting on everything you have deployed. Qwen 3.8-Max runs at its highest reasoning setting when the effort field is unset, and its highest setting was its worst performer in independent testing. A team that never touches that parameter is running the configuration that costs the most per solved task, and they will not know it until the bill arrives.
