The p99 cold start gap is the metric that defines real-world inference.

In 2026, evaluating cold start latency on GPU cloud platforms requires a focus on p99 metrics rather than the typical p50 claims.

3 min readMachine Learning

The p99 cold start gap is the metric that defines real-world inference, and the fact that no one publishes it should tell you everything you need to know. The user, u/yukiii_6, has put their finger on a problem that every serious infrastructure engineer has felt but few have articulated: p50 is a vanity number, and p99 is where your users actually live. When a request times out or a pod spins up two seconds late, it is never the median experience that gets paged. It is the tail. And the tail is where trust goes to die.

What makes this particularly damning is the silence around load. The question of whether p99 degrades under high utilization is not academic; it is the difference between a platform that works in a demo and one that survives a traffic spike. If providers are only measuring cold starts in a quiet test environment, they are measuring a fantasy. The real world is noisy, congested, and full of noisy neighbors. And if multi-provider pooling only improves p50, then it is not solving the problem it claims to solve. It is just shifting the median around while the tail stays sharp. The logic of routing to available capacity is sound, but sound logic without published data is just a hypothesis wearing a business suit.

The deeper issue is the conflation of infrastructure queue time with model loading time. These are not the same thing, and treating them as interchangeable is how marketing claims get built on quicksand. If a provider says "cold start in 500ms," what they often mean is "our scheduler is fast," not "your 70B model is loaded and warm." For someone running RTX 5090s and H200s, that distinction is not a footnote. It is the difference between a model that responds in time for a user-facing action and one that times out entirely. The industry needs a methodology that separates these variables, and until someone publishes it, every benchmark should be treated as suspect.

The practical takeaway is simple: demand better data. Ask your provider for p99 cold start under load, ask for a breakdown of queue time versus model loading, and ask how multi-provider pooling behaves when the tail is on the line. If they cannot answer, that is an answer in itself. The gap between p50 and p99 is not a technical curiosity; it is the gap between a platform that is merely functional and one that is actually reliable. And for anyone running user-facing inference, that gap is where the real cost lives.

From Machine Learning

doing infrastructure evaluation for inference workloads and running into the same problem everywhere: every platform publishes p50 cold start claims or median startup times. nobody publishes p99. and p99 is the number that shows up in support tickets and SLA violations, not p50

what I’m specifically trying to understand:

Read the original at Machine Learning