Most of us treat an LLM benchmark score like a published statistic: stable, reproducible, and true until a new version arrives. The work from AI Stupid Level challenges that assumption with data that is hard to dismiss. Across 31,352 repeated observations of 49 models, the standard deviation within a single day was 2.80 points, while the standard deviation between daily medians was 8.43 points. That is not a rounding error; it is a signal. The report is careful to note that this does not prove providers are silently swapping models, and we appreciate that restraint. But it does prove that a single score captured on a Tuesday tells you less about a model's behaviour than about the conditions under which it was measured.
The practical takeaway for anyone building on top of these APIs is uncomfortable but urgent: your evaluation pipeline is not measuring a model, it is measuring a moving target that you only partially observe. The decision to separate availability failures from valid task outcomes is particularly sharp. If a model times out or returns an error, that is not a capability signal, it is an infrastructure signal, and conflating the two corrupts your baseline. Likewise, their emphasis on versioned benchmark configurations is the kind of discipline that sounds obvious until you inherit a dashboard built on unlabelled historical runs. We would tell a reader who asks, "Should I trust this week's benchmark score?" to ask a better question: "Compared to what baseline, under what exact conditions, and with what variance?" If you cannot answer that, you are not benchmarking; you are guessing with a leaderboard.
What impresses us most is the willingness to leave the live task bank unpublished. In an era where benchmark contamination is a known disease, the author has chosen a middle path: enough transparency to inspect the methodology, enough secrecy to preserve the measurement's integrity. That is not a compromise; it is a design principle. We would push them further on one point: their question about whether daily medians are the right primary unit for change detection. Daily medians smooth away within-day variation that might be the very signal worth tracking, especially if providers route traffic to different hardware pools at different times. We would model individual repeated observations directly, with a hierarchical model that treats model version and serving environment as nested random effects. That is harder, but it is also more honest about what is actually happening under the hood.
The open question we keep circling back to is this: if a model's behaviour drifts by three times its own noise floor, what does that do to the trust you place in any single evaluation run? The work does not answer that, but it gives you the tools to ask it properly. For our readers, the concrete action is to start tracking your own evaluation data as a time series, not as a snapshot, and to demand version metadata from your providers even when they do not volunteer it. The next time someone quotes a benchmark score at you, ask to see the daily medians. Their silence will tell you more than the number ever could.