speculative decoding

Turn Idle CPU Power into Faster AI Token Generation

Speculative decoding is often framed as a GPU-only play, but DFlash proves otherwise.

4 min readTowards Data Science
Turn Idle CPU Power into Faster AI Token Generation

The numbers are hard to ignore. When DFlash delivered 3.92x the autoregressive throughput with Qwen3.5-9B on Intel Xeon 6 at concurrency 1, it wasn't just a benchmark win. It was a quiet argument that the hardware you already own might be the most underutilized asset in your inference stack. Speculative decoding on CPUs works because it turns idle compute into a betting game: a smaller model drafts, a larger model verifies, and nobody changes the output distribution. That is the kind of pragmatic innovation we like to see, the kind that improves latency without asking you to rearchitect your entire data pipeline. It reminds us of the broader lesson we explored with Clean Data Starts With Catching AI Slop Before It Skews Your Model, where small upstream choices created outsized downstream effects. Here, the upstream choice is simply deciding to use compute that would otherwise sit idle.

The acceptance metrics matter more than the raw speedup. When you hear "3.92x," your first instinct should be to ask: at what cost? The answer, according to the tests, is none in terms of output quality. DFlash preserves the model's output exactly, which means you are not trading correctness for speed. That is the difference between a clever hack and a durable technique. In our view, this is where speculative decoding earns its place alongside other practical optimizations. It is not flashy, and it does not pretend to reinvent the wheel. Instead, it asks a simple question: why let a 9B model generate tokens one at a time when a smaller draft model can do the heavy lifting and the larger model can simply approve the work? That division of labor is elegant because it is understandable. You do not need a PhD to grasp why verification is faster than generation, and you do not need a GPU cluster to benefit from it.

If you are building systems that rely on CPU inference, this is worth a hard look. The concurrency 1 result is particularly telling because it isolates the technique's core value without the noise of batch effects. At higher concurrency, you might see different behavior, but for low-latency, single-stream requests, this is a meaningful win. It also suggests that the gap between CPU and GPU inference is not as wide as the marketing departments want you to believe. That is a refreshing counterpoint to the constant pressure to buy more expensive hardware. We would tell any reader who is skeptical to run the test themselves, but start with the acceptance rate. If the draft model is accepted even 70% of the time, you are already ahead. And if you are curious about how other optimizations compare, our piece on Exploring Real-World Computer Vision: Deployments, Edge Models, and Current Challenges shows that similar constraints apply when you are trying to squeeze performance out of limited resources.

The open question is whether this scales beyond the specific configuration tested. Intel Xeon 6 is a capable platform, but the real test will come when speculative decoding meets messier, more varied workloads. The technique's reliance on a draft model means its effectiveness depends on how well that draft model aligns with the target distribution. When they diverge, the acceptance rate drops and the speedup shrinks. That is the detail to watch. For now, the takeaway is clear: you can get nearly 4x faster token generation on CPUs without changing your model's output. That is not a promise of a revolution. It is an invitation to explore what your existing infrastructure can do when you stop assuming it has to be slow.

From Towards Data Science

Speculative decoding can turn underused CPU compute into faster token generation, without changing the model's output. In our vLLM tests, DFlash delivered 3.92x the autoregressive throughput with Qwen3.5-9B on Intel Xeon 6 at concurrency 1. We break down where the speedup comes from, explain the acceptance metrics, and show what determines whether speculation pays off.

The post Speculative Decoding on CPUs: Nearly 4x Faster Token Generation with DFlash appeared first on Towards Data Science.

Read the original at Towards Data Science