When a model's accuracy drops by more than twenty points, most teams scramble to retrain or wait for a new release. That is the wrong instinct. The real insight from this benchmark is that the smartest spreadsheets, the ones that care about correct answers, don't rely on the model alone. They let the model write its own code, then execute that code against the data. That is not a workaround. It is a better architecture.
The numbers are clear. GPT-5.4-mini's vanilla accuracy fell from 69.5% to 47.2% across twelve tasks. The official RLM implementation dropped nearly as far, from 69.7% to 50.2%. But the implementation that forces the model to generate Python and run it in a REPL held at 69.5%, a drop of only 3.2 points. On AIME 2025, the gap is even starker: 80% accuracy versus zero. Zero. The bare model guesses; the code-writing model computes. That is the difference between a tool that fails under pressure and one that adapts.
What this means for you is practical. If your spreadsheet or data tool depends on the model's raw pattern-matching to answer every query, you are vulnerable to every update, every model change, every terse output. The architecture described here absorbs those changes. It does not ask the model to be smarter. It asks the model to be a programmer, then runs the program. That approach also delivers 5.1 times fewer tokens than the official RLM method, at 3.2 times lower cost. And it works with every model, not just this one.
You should evaluate your own workflows. Where are you asking the model to attend to all the data and guess the answer? Where could you instead let it write a query and execute it? The evidence from this test is not theoretical. It is 1,800 evaluations across twelve tasks, repeated across two model versions. When the model gets worse, the code-writing architecture stays stable. That is the kind of resilience worth adopting, not because it is innovative, but because it is more reliable.