formula generator

Hold AI Accountable: A New Benchmark Tests Physics Without Guesswork

Introducing a novel benchmark designed to evaluate the performance of large language models (LLMs) against fundamental physics laws.

3 min readMachine Learning

This benchmark is exactly the kind of accountability the AI industry has been avoiding. A developer got tired of watching large language models produce confident nonsense in physics, so he built a test that cuts through the performance theater. No LLM-as-judge, no vibes, just symbolic math. That matters because the gap between a model that *sounds* right and a model that *is* right is where real-world mistakes happen, and right now, that gap is wide enough to drive a truck through.

The results are instructive. Look at the Gemini lineup: the flash-image variant scored 88.6 percent, while the pro model, the one you'd expect to be the smartest, landed at 22.1 percent. That is not a rounding error. The pro model kept falling for the same basic trap: forgetting the ½ in kinetic energy. It also bombed gravitational force questions. Meanwhile, the flash-lite preview scored 72.9 percent. The pattern is clear: bigger, more expensive models are not necessarily more reliable. They are just more confident when they are wrong.

For anyone who uses AI to do real work, this is not an academic curiosity. If you are a student, an engineer, or a researcher who relies on these tools to check your physics, the benchmark reveals a dangerous blind spot. Anchoring bias, unit confusion, formula traps, these are not exotic edge cases. They are the kind of mistakes a first-year student learns to avoid. Yet the models stumble on them consistently. Bernoulli's Equation crushed every model tested, including the best one, because pressure unit confusion between pascals and atmospheres is apparently a universal weakness. If your workflow touches fluid dynamics, you have been warned.

The practical takeaway is simple: treat every AI-generated physics answer as a first draft, not a final result. The developer's approach, procedural question generation, symbolic math grading, no human judgment involved, is the right way to test these systems. It is reproducible, transparent, and unforgiving. We hope more people follow his lead, test more models, and publish the results. Until the industry starts owning its errors instead of papering over them with marketing, benchmarks like this are the only honest feedback loop we have.

From Machine Learning

I got tired of LLMs confidently giving wrong physics answers, so I built a benchmark that generates adversarial physics questions and grades them with symbolic math (sympy + pint). No LLM-as-judge, no vibes, just math.

The benchmark covers 28 physics laws (Ohm's, Newton's, Ideal Gas, Coulomb's, etc.) and each question has a trap baked in:

Read the original at Machine Learning