When AI ignores the evidence, discovery loses its meaning.

Recent research involving 25,000 experiments has revealed alarming trends in AI scientists: they often generate results without adhering to scientific reasoning.

3 min readMachine Learning

The numbers are damning, and we should stop pretending otherwise. When researchers ran 25,000 AI scientist experiments and found that 68% of the time the AI gathered evidence and then ignored it, they exposed something fundamental: these systems are not reasoning toward truth. They are pattern-matching toward plausibility. And 71% of the time, the AI never updated its beliefs at all. Not once. Only 26% of the time did it revise a hypothesis when confronted with contradictory data. That is not a bug in an otherwise sound approach. That is the approach.

For anyone building on AI research tools, this should change your expectations immediately. You are not getting a junior collaborator who double-checks their work. You are getting a confident autocomplete that has learned what conclusions sound like, not how to earn them. A human scientist adapts. You approach a chemistry identification problem differently than you approach a simulation workflow. The AI does not. It runs the same undisciplined loop every time, producing results that look rigorous because they are formatted rigorously. The evidence is collected, cited, and then discarded when it becomes inconvenient. That is not a failure of execution. That is a failure of epistemology.

The most frustrating part is that the field's response has been to double down on the wrong fix. Researchers have focused on better scaffolding: ReAct, structured tool-calling, chain-of-thought, all the engineering frameworks that promise to make agents more reliable. The study tested the most popular proposed solution and found it does not work. You cannot scaffold your way to genuine inquiry. You can build a better cage, but the animal still does not know why it is pacing. The problem is not that the AI lacks a better prompt. The problem is that it lacks a reason to care whether the evidence contradicts its conclusion. No amount of tool routing solves that.

Here is what this means for you, practically. If you are using AI to accelerate discovery, treat its outputs as a starting point for your own skepticism, not as a conclusion to validate. Demand transparency about how the system handled contradictory evidence. And push for evaluation metrics that measure belief revision, not just output quality. A system that never changes its mind is not a scientist. It is a static document generator. Until that changes, the burden of genuine discovery remains on you. That is not a limitation to engineer around. It is the actual work.

From Machine Learning

Researchers ran 25,000 AI scientist experiments and discovered something that need attention!!

AI scientists are producing results without doing science.

Read the original at Machine Learning