ARC-AGI-3

How small AI models closed the gap on a test built for humans

For the first time, small local AI models have matched average human performance on a benchmark built specifically to prove human superiority, and it happened in just the last 30 days.

3 min readMachine Learning
How small AI models closed the gap on a test built for humans
Top ARC-ΑGI-3 scores on Kaggle just went from 7% to 56% [N]

The gap between small AI models and human performance just collapsed on a benchmark built to prove we were irreplaceable. Over the last 30 days, top scores on the ARC-AGI-3 Kaggle leaderboard jumped from 7% to 56%, with local models, the kind any competitor can run on their own hardware, matching average human results on a test explicitly designed to highlight human reasoning. This is not a slow creep. It is a sudden, measurable shift that changes what we should expect from the tools we use every day.

For anyone who has felt the tension between wanting more intelligent software and worrying about complexity, this moment matters. Small models are not the exclusive domain of massive data centers. They run in harnesses on consumer machines, and they are now solving puzzles that were supposed to separate us from them. That has direct implications for how we think about productivity tools. If a modest local model can close a reasoning gap in 30 days, the distinction between a "dumb" spreadsheet and an "AI-native" one becomes less about futuristic promises and more about practical, present capability. We have seen this pattern before in our coverage of how Opus 5.5 redefines what a spreadsheet benchmark should look like, where the measure of a tool shifted from feature counts to actual reasoning outcomes. The ARC-AGI-3 results are another data point in that same trajectory: performance on human-centered tests is becoming a baseline, not an aspiration.

We should be clear about what this does not mean. It does not mean small models have achieved general intelligence or that humans are obsolete. The benchmark is narrow by design. But it does mean that the assumption "AI cannot handle nuanced, pattern-based reasoning" is no longer defensible for even lightweight systems. That changes the conversation around adoption. When we built IQRAX to give AI models both freedom and verifiable authority, we argued that the real barrier was trust, not capability. These Kaggle results reinforce that argument: the reasoning capability is here; the question is how we integrate it without losing control. Similarly, the work on OpenAPPA and its perfect security record against prompt injection shows that the infrastructure to deploy these models safely is evolving alongside their raw performance.

The specific takeaway is this: if you have been waiting for AI to be both capable and accessible enough to handle your actual workflows, the window just narrowed significantly. The 56% score on ARC-AGI-3 is not a ceiling; it is a floor that rose in a single month. Watch what happens in the next 30 days.

From Machine Learning

https://preview.redd.it/gkjgii48gfth1.png?width=575&format=png&auto=webp&s=4e1c35bdacf41d18a6409baedcddc9432ec79814

This happened over the past 30 days. So smallish local models (Kagglers can only use those), in a harness, just started beating average humans at a benchmark intentionally designed to show human superiority.

Read the original at Machine Learning