Scaling Qwen 3.5 to a million tokens per second with practical findings

In this detailed analysis, we explore the impressive achievement of processing 1.1 million tokens per second with the Qwen 3.5 27B model on 96 B200 GPUs, utilizing vLLM v0.18.0. The findings reveal that using DP=8…

3 min readMachine Learning

The team at Google Cloud just posted a real-world benchmark that tells us more about inference at scale than a dozen press releases. They pushed Qwen 3.5 27B to 1.1 million tokens per second across 96 B200 GPUs, and the results are worth studying closely. What matters most here is not the raw number, but the practical decisions that made it possible, and the ones that didn't work.

The most instructive finding is that data parallelism nearly quadrupled throughput over tensor parallelism for this model. That is a concrete signal for anyone running a model of similar size: on B200s, the model is too small for tensor parallelism to help. Multi-Token Prediction was even more decisive. Without MTP-1, GPU utilization sat at zero percent. With it, the system came alive. But MTP-5 crashed with a CUDA error, so the benefit has a ceiling. These are the kinds of details that save teams weeks of trial and error.

The scaling efficiency numbers are strong: 97.1 percent at eight nodes, 96.5 percent at twelve. Time per output token stayed flat at roughly 46 milliseconds regardless of node count. That is the behavior you want when scaling out, predictable latency that does not degrade as you add hardware. The Inference Gateway, which routes requests based on KV-cache awareness, added about 35 percent overhead compared to a simple round-robin approach. That overhead is a tradeoff worth understanding: smarter routing costs cycles, and whether the gain in cache reuse justifies the latency hit depends entirely on your workload mix.

This benchmark was run with worst-case assumptions: no prefix cache hits, input length of 1024 tokens, output length of 512 tokens. The results are not cherry-picked. They represent a floor, not a ceiling. For anyone building production inference infrastructure, the takeaway is straightforward: start with data parallelism, invest in Multi-Token Prediction up to a stable point, and measure your routing overhead before assuming smarter is faster. The path to a million tokens per second is paved with decisions that are specific to your model and your hardware, not with generic optimizations.

From Machine Learning

Wrote up the process of pushing Qwen 3.5 27B (dense, FP8) to 1.1M total tok/s on 96 B200 GPUs with vLLM v0.18.0.

InferenceMAX methodology, input-len=1024, output-len=512, 0% prefix cache hit. Worst-case numbers.

Read the original at Machine Learning