Measure your coding agent against a new open-source benchmark suite.

Introducing an open-source benchmark suite designed to evaluate coding-agent performance in real-world software tasks.

3 min readMachine Learning
Measure your coding agent against a new open-source benchmark suite.
Open-source 9-task benchmark for coding-agent retrieval augmentation. Per-task deltas +0.010 to +0.320, all evals reproducible [P]

A benchmark that measures whether a coding agent performs better when it can search academic literature before writing code is exactly the kind of concrete evidence this field needs. The results are clear: across nine everyday software tasks, the agent with access to retrieval outperformed the baseline in every single case, with deltas ranging from a modest 0.010 to a striking 0.320. This is not a theoretical argument about the future of AI-assisted development. This is a reproducible finding, published openly, with every prompt, evaluation script, and prediction file available for inspection.

What makes this benchmark valuable is its honesty about the practical reality of coding agents. The author deliberately chose tasks that mirror actual engineering workflows, test generation, text-to-SQL, PDF extraction, contract analysis, code review, classification, prompt selection, LLM routing, and summarization evaluation. These are not exotic ML challenges. They are the kinds of problems that developers actually bring to agents today. The selection criteria were refreshingly pragmatic: each task has an unambiguous quantitative metric, a baseline performance well below ceiling, standard datasets where available, and an evaluation that runs on a free API key in roughly ten minutes. This is a benchmark built for everyday use, not for conference leaderboards.

Some of the largest gains came from techniques published after the agent's training cutoff. The contract extraction improvement came from BEAVER and PAVE, both 2026 papers. The test generation improvement came from mutation-aware prompting strategies that enumerate AST-level mutations. The agent could not have reached these approaches from parametric memory alone. Ten of the fifteen most-cited sources across all experiments were published in 2025 or later. This is the conservative argument for retrieval: it is not about making agents smarter, but about making them aware of knowledge they fundamentally do not possess. Retrieval as a memory expansion, not an intelligence boost.

The author also documents the failures, which is rare and welcome. Self-refinement hurt text-to-SQL performance because the agent second-guessed correct queries after reading academic work on SQL ambiguity. Two retrieved techniques in the auto-research experiment were architecture-incompatible and had to be discarded. Retrieval surfaces better options, but it does not guarantee wins. The agent still needs to evaluate, filter, and decide. The benchmark makes this visible. If you are building or evaluating coding agents, this repo offers a replicable methodology for measuring whether your system benefits from the same kind of tool access. The data is in the repository. The methodology is documented per task. The opinion here is plain: this is how benchmarks should be built.

From Machine Learning

Sharing an open-source benchmark suite (paper-lantern-challenges) that measures coding-agent performance with vs without retrieval-augmented technique selection across 9 everyday software tasks. Disclosure: I'm the author of the retrieval system under test (paperlantern.ai/code); the artifact being shared here is the benchmark suite itself, not the product. Every prompt, agent code path, and prediction file is in the repo and reproducible.

Read the original at Machine Learning