The most quietly subversive finding in worldproof: diagnosing where world-model predictions break and a measurement of when pixel metrics stop being able to rank models at all is not that pixel metrics fail. It is that they fail in a specific, measurable, and fixable way, and that the fix has nothing to do with building a better model. The author ran a "predict nothing changes" baseline against real robot footage and watched it score a 0.983 SSIM and 53.9 dB PSNR on a 30fps recording, with error that refused to grow across a six-step horizon. That is not a broken metric. That is an evaluation setup with zero discriminative power, and it is far easier to miss because the numbers still look impressive.
What makes this worth pausing over is the shape of the failure. The same baseline on DROID footage, extended to 48 steps, reveals three regimes: a dead zone at the start where everything ties, a steep monotonic decline where models are actually separable, and a floor around 0.20 SSIM where prediction has fully decorrelated and everything ties again at the bottom. The usable window for this kind of footage sits somewhere between 8 and 24 steps, and this is a property of frame rate times task speed, not a universal constant. That is why the Explore How AI World Models Are Empowering the Next Generation of Robotics conversation needs to move past "better architectures" and toward "better evals." You cannot rank what you cannot separate, and if your horizon is too short or too long, you are not measuring model quality, you are measuring the frame rate.
The practical takeaway for anyone building or evaluating world models is uncomfortable: your summary scalar is lying to you. The author found that including step 0 inflates every averaged number, because a copy baseline gets a nearly free first step whenever the frame rate is high relative to scene motion. On the 30fps recording, step 0 scores 119.8 dB, dragging the horizon-averaged scalar from 32 up to 53.9. That means a model can appear twice as good as it is, purely because the evaluation rewards frame rate over prediction quality. Curves are the honest thing to report, and the scalar definition is treated as an open problem, which is exactly the right instinct. If you are comparing models, do not look at the headline number. Look at the curve, identify your own dead zones, and measure your usable window on your own data.
The author also flags that the n=8 version of the SO-101 run gave intervals wide enough to overlap DROID completely, and only the jump to n=64 produced trustworthy separation. That is a reminder that small rollouts are not just noisy, they are actively misleading. And the LPIPS inconsistency, where the metric points the wrong way on the masked variant with no clean explanation, is a loose end that deserves attention. What we would tell a reader who asks us about this: do not inherit a horizon from a paper that used different data. Run the copy baseline, plot the curve, and find the stretch where your own task actually separates models. The tool is open-source and runs on a laptop, so there is no excuse to skip this step. The specific thing to watch next is whether the community adopts horizon curves as a standard practice, or whether the allure of a single scalar wins again. Our bet is on the latter, and that is precisely why this work matters.