When LLMs consult research papers, automated experiments gain a 3.2% edge.

In a controlled experiment, a Claude Code agent demonstrated a 3.2% improvement in results by accessing over 2 million computer science papers during automated hyperparameter searches. Utilizing Karpathy's autoresearch…

3 min readMachine Learning
When LLMs consult research papers, automated experiments gain a 3.2% edge.
[R] Controlled experiment: giving an LLM agent access to CS papers during automated hyperparameter search improves results by 3.2%

This experiment is the most practical argument we've seen for giving AI agents live access to research. A 3.2% improvement on a problem as well-trodden as TinyStories is not a fluke; it is a signal that the ceiling for automated experimentation is far higher than we assume. The agent without papers was limited to techniques already memorized in its weights. The paper-augmented agent found AdaGC, a method published after its training cutoff, and correctly applied the sqrt batch scaling rule on its first try. That is the difference between an agent that guesses and an agent that reads.

The practical takeaway for anyone running automated experiments is clear. Your LLM agent has a wall of knowledge it cannot see. It may have encountered a technique during training, but unless you prompt it to search, that knowledge stays latent. The paper-augmented agent didn't just find more techniques; it found the *right* ones faster. When the agent without papers halved the batch size and diverged because it didn't adjust the learning rate, the paper-augmented agent retrieved the scaling rule and applied it correctly. That saved time, compute, and the frustration of a failed run.

We also appreciate the honesty in the limitations. Some of the gain may come from the agent spending more time reasoning about each technique, not just from the paper content. That is still a win. Forcing the agent to search, read, and synthesize introduces a deliberation step that improves outcomes regardless of the source. The fact that DyT and SeeDNorm were tried and reverted shows the agent was not blindly applying every paper it found. It tested, observed, and discarded. That is actual scientific reasoning.

The next step is to replicate this at larger scale and on less-explored problems. TinyStories was a deliberate stress test to make the comparison harder. If the gap widens on novel domains, we will see a fundamental shift in how automated research works. For now, the message is simple: if you are running automated experiments without giving your agent access to the literature, you are leaving a measurable advantage on the table. Build the search layer. The 3.2% edge is just the start.

From Machine Learning

Ran a controlled experiment measuring whether LLM coding agents benefit from access to research literature during automated experimentation.

Two identical runs using Karpathy's autoresearch framework. Claude Code agent optimizing a ~7M param GPT-2 on TinyStories. M4 Pro, 100 experiments each, same seed config. Only variable — one agent had access to an MCP server that does full-text search over 2M+ CS papers and returns synthesized methods with citations.

Read the original at Machine Learning