If you feel constrained by traditional benchmarks that measure rote memorization, the Nonobench results offer a far more honest assessment of where AI logic actually stands. A set of 49 large language models, tested on nonogram puzzles from 5x5 to 20x20 grids, reveals a steep and telling decline: solve rates drop from 85% on the smallest puzzles to just 20% on 15x15 grids, and on the hardest 20x20 puzzles, only two models managed to solve more than half. This isn't a story about which lab won, it's a story about how quickly even the most advanced models hit a wall when the task demands genuine step-by-step reasoning rather than pattern matching. For anyone who has watched How small AI models closed the gap on a test built for humans, this feels like a sobering sequel: where smaller models once surprised us by catching up, here the gap between human-like logic and machine processing remains stubbornly wide.
The Nonobench methodology is refreshingly strict: each model receives the row and column clues once, returns the full grid in a single attempt, and gets no tools, no iterative feedback, no second chances. That one-shot constraint is precisely what makes the results meaningful. When a model fails a 20x20 puzzle that has a unique solution, it is not because the puzzle is ambiguous, it is because the model lost track of the constraints before the logic got hard. The researchers observed that when the clues were presented as a single 400-character string, most models simply lost count. They had to restructure the input into an array of 20 row strings just to get usable results. This is a concrete, measurable weakness: these systems cannot reliably hold and manipulate a moderate amount of structured information in a single pass. It echoes the lessons from OpenAPPA Turns the Tables on Prompt Injection With a Perfect Security Record, where security required explicit structural guards, here, logic itself requires structural scaffolding that the models cannot provide on their own.
The practical takeaway for anyone building workflows around AI is straightforward: do not assume that a model that can write a sonnet or pass a bar exam can reliably solve a puzzle that a patient human can crack with a pencil and grid paper. The Hard mode puzzles were deliberately chosen to include five that cannot be solved by line logic alone, they require looking at the whole board, holding multiple possibilities in mind, and backtracking when a contradiction appears. That is the kind of reasoning that remains, for now, distinctly human. The best model, Claude Opus 5.5, solved 8 of 10 Hard puzzles, which is impressive until you consider that 11 of the 15 models tested on that mode solved none. The gap between the leader and the field is wider than the gap between the leader and a competent human solver.
What we are watching is a boundary being drawn in real time. The Nonobench code is open-source, the puzzles are public, and the methodology is clear enough that anyone can replicate it. That transparency is valuable because it gives us a specific, reproducible failure mode to track. The question to watch is not whether next year's model will score higher, it almost certainly will, but whether the improvement comes from better architecture or simply from training on more puzzle data. If the former, we are making real progress toward reasoning. If the latter, we are just watching the boundary of memorization expand. The next time you see a benchmark claiming near-perfect scores, ask yourself whether the test allows the model to cheat by pattern-matching its way to a plausible answer. Nonobench does not, and that is precisely why its results matter.
