The International Mathematical Olympiad has long been the gold standard for measuring human mathematical reasoning, so it is fitting that it has become a proving ground for large language models. The recent comparison of frontier models, open-weight systems, and custom harnesses on the IMO 2026 problems offers a refreshingly sober look at where we actually stand. The results are not about hype; they are about capability, and the gap between what a model can do alone and what it can do with orchestration is the story we should be paying attention to. What stands out is the role of the harness. The data shows that for models like Claude Sonnet and Opus, the webapp performance was poor, but the gap narrowed significantly with provider-specific tools like Claude Code, and improved even further with AutoFyn, a customizable multi-agent harness. This is a practical insight for anyone building on top of LLMs. The model is only one part of the equation; the scaffolding around it can be the difference between a mediocre result and a strong one. Yet the same data reveals a hard ceiling: even with a 20-hour run, the key reduction on the hardest problem, P3, was missed by every sub-frontier model in every harness. The harness supplied retrieval and verification, but it did not supply the creative leap. That is a humbling reminder that while we can optimize workflows and improve retrieval, the core reasoning step remains the bottleneck. The fact that the grading was done by a former IMO medalist, with manual verification, adds a layer of credibility that is often missing in AI benchmarks. It is not just a matter of automated scoring; a human who knows what a correct solution looks like confirmed the results. And the hallucination issue, such as the false solution on P3 by Sonnet, persists even in a verifiable domain like mathematics. This is not a minor flaw; it is a fundamental challenge. If a model can confidently produce a wrong proof in a field where correctness is binary, then we need to be cautious about any application where errors are not immediately visible. This connects to the broader conversation about practical decision-making, as explored in our piece on Jev vs LLMs: Evaluating AI for Practical Decision-Making, where the stakes are lower but the pattern is the same: confidence without guaranteed accuracy. For our readers, the takeaway is not that frontier models are unbeatable, but that the distance between a capable model and a reliable system is filled with engineering. The open-weight model GLM performed roughly at the same level as Sonnet without a harness, and improved similarly with AutoFyn. That is significant because it suggests that the gap between proprietary and open models is not as wide as some might assume, especially when you factor in the ability to customize the orchestration layer. This is a direct challenge to the assumption that only the largest labs can build competitive AI systems. It also aligns with the ongoing work in distributed training and inference, where the focus is on making the underlying infrastructure more accessible, as discussed in Unlock LLM Training: A Practical Guide to Distributed Algorithms. The future is not just about bigger models; it is about smarter ways to use the models we have. The real question this raises is whether we are approaching the limit of what harness engineering can achieve. The frontier models, Sol and Fable, scored perfectly or nearly so regardless of the harness. That suggests that for the most advanced systems, the model itself is the primary driver of performance, and the harness is just a support structure. But for everyone else, the harness is where the gains are to be found. The practical consequence for you is this: if you are building an AI-powered tool, do not assume that the model is the ceiling. Invest in the orchestration, the verification, and the retrieval, because that is where you will see the most immediate improvement. And if you are relying on a model to reason through a complex problem, build in a verification step, because even in a field as precise as mathematics, hallucination remains a real risk.
LLMs
How AI masters new math problems reveals its true reasoning potential
The IMO 2026 benchmark results reveal a clear hierarchy: frontier models like sol and fable handle the math olympiad with near-perfect scores, while others, including sonnet and opus, lag without a harness.
4 min readMachine Learning

There are a few reasons why problems from International Mathematical Olympiad function as a good benchmark for LLMs:
- The problems are new, not included in the training data of any model