Most teams treat WAN latency as a hard wall for distributed inference. The math feels unforgiving: if every token needs a round trip, then 86 milliseconds of public internet RTT caps you at a handful of tokens per second. That wall is exactly what the ShardFlow project, built by u/katua_bkl, just pushed through. By splitting Qwen2.5-7B across two T4 nodes in separate GCP regions, with an AWS relay in between, they turned a naive 4.92 TPS baseline into 28.10 TPS peak. Not by shrinking the distance, but by changing what the distance costs. That distinction is worth pausing on, because it reframes how we should think about distributed AI altogether.
The core trick is neural speculative decoding, where a small draft model generates K tokens locally, and the larger model verifies them in one pass. In a single-machine setup, speculation is a nice efficiency win. Across a public WAN, it becomes existential. The key insight: latency stops being a per-token tax and becomes a per-round cost. With K=8, you commit roughly four tokens per round trip instead of one. At 86ms RTT, that is the difference between a toy demo and a usable system. The numbers back it up. But the more instructive story is the CUDA Graphs discovery. The draft generation was launching around 1,500 kernels per round from a Python loop, each paying 8 to 10 microseconds in overhead while the GPU sat idle 65% of the time. Capturing the full 0.5B forward pass as a single graph and replaying it with one driver call dropped draft latency from 112ms to 25ms. That is a 4.5x improvement from removing software overhead, not from touching the model or the network.
What does this mean for you, practically? It means the barrier to distributed inference is not silicon or bandwidth. It is orchestration discipline. This echoes lessons from Unlock LLM Training: A Practical Guide to Distributed Algorithms, where the fundamentals of sharding and pipeline parallelism matter more than any single hardware spec. And it aligns with the resource-constrained engineering approach that Build Scalable Products with Less: Engineering Lessons from Startups advocates: you do not need the newest GPU to get serious throughput, you need to eliminate inefficiency. The v2.1 fix is a masterclass in that mindset. Most engineers would have blamed the network or the model size. They found the real culprit in their own toolchain.
There is a bigger question here, and it is not about TPS. ShardFlow is open source, and the project is transparent about the stack: zero-copy Rust relay, StaticCache with in-place KV rewind, and meta-device slicing to avoid loading 15GB into CPU RAM. That is a lot of moving parts, and it works. But it also suggests that the next frontier is not just making distributed inference faster, it is making it boring enough to adopt. The fact that they hit 14.43 TPS on Qwen2.5-14B with NF4 quant across the same two nodes is a compelling data point that smaller clusters can punch well above their weight. The specific thing to watch is whether the CUDA Graphs approach becomes a standard pattern in inference engines, because if a single driver call can replace 1,500 kernel launches, then every framework that ignores this is leaving a 4x speedup on the table. That is not a marginal gain. That is a new baseline for what we should expect from distributed inference.