The scoreboard is lying to us. When an AutoResearch-style agent takes a human-defined problem, a human-built evaluator, and a human-selected starting point, then climbs toward a better number, we are measuring optimization, not discovery. That distinction matters, and it matters more as these systems get faster. The user behind this project, /u/Only-Aardvark2568, has put a finger on the exact tension: a search that explores thousands of variants is genuinely useful, but it is not the same as asking whether the problem itself should be re-framed. We would argue that score improvement is the easy part. The hard part, the part that looks like actual research judgment, is knowing when the metric is the wrong target.
This is not a semantic quibble. It has practical consequences for how we build and trust these tools. If we conflate a high score with scientific insight, we risk optimizing for local peaks while missing the transferable principle sitting one valley over. A researcher asks whether a result generalizes, whether the formulation itself is flawed, whether a different direction is more promising. An iterative loop rarely asks those questions because it is not rewarded for asking them. Consider how Reward-aware search replaces blind sampling in best-of-N generation refines the sampling process itself: that is a smarter way to search a known space, not a way to question whether the space is worth searching. The same logic applies here. Better optimization within a human-shaped arena is valuable, but it is a tool, not a research agenda.
What would it take to move beyond that? The agent needs a mechanism to challenge its own constraints. That means building evaluators that can be questioned, not just maximized. It means rewarding agents for proposing alternative problem formulations, for testing transfer across tasks, for identifying when the current objective is leading to a dead end. None of this is automatic. It requires a deliberate design choice to treat the human-defined setup as a hypothesis rather than a given. This connects to a broader pattern we are seeing across the field: the shift from predicting outcomes to modeling the world itself, as in Predict human behavior with a world model built from scratch, which implies a much richer notion of what an agent should understand before it acts. And it touches on governance, because if we cannot tell the difference between optimized search and genuine insight, we will make poor decisions about where to deploy these systems, much like the legal questions raised in Court rules warrantless Flock plate searches breach Fourth Amendment rights, where the tool's convenience outpaced the scrutiny of its purpose.
The concrete takeaway is this: if you are building or funding an AutoResearch system, add a second metric that has nothing to do with the task score. Measure how often the agent proposes a change to the problem definition, how often it abandons the given direction for a better one, and how often its results transfer to an unseen task. Those numbers will tell you whether you have a search engine or a researcher. Until then, treat every impressive score improvement as a strong local result, not a scientific breakthrough. The distinction is not academic. It determines whether these systems expand our understanding or just polish the assumptions we already made.