multi-modal large language models

AI models stumble when spreadsheet rules demand creative combination

In 2024, MIT and Virginia Tech researchers set three state-of-the-art multimodal models against the deceptively simple puzzles of Baba Is You, requiring creative rule manipulation.

3 min readMachine Learning

The gap between what our most powerful AI models can achieve and what they actually understand has never been more visible than in the recent testing of GPT-4o, Gemini-1.5-Pro, and Gemini-1.5-Flash against the puzzle game Baba Is You. These are models that can draft legal briefs, generate working code, and hold nuanced conversations. Yet they fail dramatically when asked to manipulate and combine simple game rules. That is not a minor engineering hiccup. That is a fundamental signal about how these systems learn, and it deserves far more attention than it received when this research was presented at ICML in 2024.

For anyone building workflows around agentic systems, this is the practical warning. The paper from MIT and Virginia Tech researchers shows that when a task requires generalization beyond pattern matching, the models break. Not because the puzzles are hard for humans, but because the rules must be actively bent and recombined. We have seen similar dynamics in our own reporting on why a single successful agent run does not mean the database agrees, where isolated wins mask deeper reliability gaps. And the lesson from mapping networks is equally relevant: compact latent representations can outperform brute-force parameter counts, suggesting that scale alone is not the answer to reasoning deficits. The connection is uncomfortable but clear: our tools are getting bigger, not necessarily smarter.

By October 2026, the landscape has shifted dramatically. We are told agentic swarms powered by tera-parameter models can ace ARC-AGI-3, solve FrontierMath tier 4 problems, and even crack Navier-Stokes theorems. On paper, this looks like the singularity arrived quietly. Yet the small key-door puzzles from this 2024 paper remain a plausible stumbling block. If these simple rule-combination tasks still trip up the latest models, then the entire narrative of exponential capability growth needs a serious asterisk. It is not enough to solve harder math problems or write better essays. The ability to manipulate abstract rules in novel combinations is a core component of what we call reasoning, and if that remains broken, then every downstream application built on these models inherits that fragility.

The authors suggest this benchmark could be a candidate for ARC-AGI-4, and they are right to make that case. Francois Chollet and the ARC Foundation should be paying close attention. But the more immediate question is for the engineers and product teams building on these systems today. Can an agentic swarm solve the Baba Is You puzzles? That is not a theoretical exercise. It is a testable, concrete benchmark that would tell us whether the gap has narrowed or whether we have simply gotten better at hiding it. Until someone runs that experiment and publishes the results, the prudent assumption is that the weakness persists. And if it does, every automation pipeline that relies on creative rule combination is running on borrowed confidence. Watch for that replication study. That is where the truth will surface.

From Machine Learning

We test three state-of-the-art multi-modal large language models (GPT-4o, Gemini-1.5-Pro, Gemini-1.5-Flash) and find that they fail dramatically when generalization requires that the rules of the game must be manipulated and combined.

The catch here is that this statement was written in a paper in 2024 that was brought to ICML conference that year. (the conference was held in Austria)

Read the original at Machine Learning