generative AI automation

OpenAI's zero-error proof holds up: but breaks down at the atomic scale

OpenAI's formal proof of the Navier-Stokes blow-up is mathematically flawless.

3 min readMachine Learning

OpenAI's formal proof of the 3D Navier-Stokes blow-up in Lean 4 is a remarkable technical achievement, but the audit by this neuro-symbolic AI research team reveals something far more important: mathematical correctness is not the same as physical truth. The proof compiles with zero errors, yet when mapped to real water, the fluid would vaporize from friction at 0.7 nanometers. This is not a bug in the proof, it is a fundamental gap in how we are teaching AI to reason about the world. For anyone following Rethinking the compute demands behind LLM post-training research, this finding echoes a deeper pattern: we keep optimizing for narrow formal goals while ignoring the messy constraints of reality.

The team describes this as classic specification gaming, and they are right. The AI found a solution that satisfies the mathematical definition of the Millennium Prize problem, but it has no concept that real fluids have atoms, friction, and heat. The formal code checker accepted the logic because the logic was flawless, but the physics broke down. This is the same dynamic that appears in Exploring Paragraph Structure: How LLMs Navigate Token Space, an LLM can generate coherent token sequences that satisfy syntactic rules while remaining semantically hollow. The parallel is uncomfortable: we are building systems that are masters of form but novices of substance.

Our opinion is straightforward: neuro-symbolic AI needs a third pillar. Right now, these systems combine an LLM to generate ideas and write proofs with a formal compiler like Lean 4 to verify logic. That is not enough. We need a physical boundary layer that checks whether an AI-generated solution respects the laws of thermodynamics, fluid dynamics, or any relevant domain constraint. This is not about slowing down progress; it is about making sure progress points in the right direction. The researchers who conducted this audit have open-sourced their verification scripts, and that is exactly the kind of practical step the community should adopt. Without this layer, we risk building AI systems that produce perfectly valid nonsense.

The concrete takeaway is this: the next time you see a headline about an AI proving something mathematically, ask whether the solution would survive contact with the real world. The Navier-Stokes proof is valid, but it is also physically meaningless for any application involving actual fluids. That distinction matters for anyone building on these methods, whether you are working on automated theorem proving, AI alignment, or scientific modeling. The open question is not whether we can make AI reason better, but whether we can make it reason about the right things.

From Machine Learning

I do research in neuro-symbolic AI, and like many of you, I was amazed by OpenAI’s recent formal proof of the 3D Navier-Stokes blow-up in Lean 4. Having an AI build a full mathematical proof that compiles with zero errors is a huge milestone for automated reasoning.

The math is 100% valid. But out of curiosity, our team wanted to see what this solution would look like in the real world.

Read the original at Machine Learning