The real bottleneck here was never S3 egress fees, and it was never Cloudflare R2's TTFB. It's the assumption that your data lake and your training cluster should be treated as two separate problems. The user in this story is doing exactly what most teams do: they hit a cost wall, they swap one object store for another, and then they're surprised when the new store's latency profile doesn't match the workload. Of course it doesn't. R2 solves egress pricing, but it doesn't solve the fundamental mismatch between a network filesystem's semantics and a GPU's need for a steady, predictable stream of bytes. The 20% idle time isn't a Cloudflare issue. It's the cost of treating object storage as if it were local NVMe.
What this means for you is that the choice isn't between paying AWS egress or building a custom cache layer. Those are both workarounds, not solutions. The custom NVMe cache path is especially seductive because it feels like engineering control, but it's really just another system to operate, another failure domain, and another reason your data pipeline becomes a full-time job. The actual question is whether you can find a zero-egress store that's designed for high-speed streaming, not just for cheap storage. That means looking for something that can sustain high throughput per connection, not just high aggregate throughput across many parallel requests. R2's inconsistent time-to-first-byte is a symptom of a system built for durability and scale, not for low-latency reads. Your data loader doesn't need a filesystem. It needs a pipe that fills up faster than the GPU drains it.
The practical takeaway is this: don't optimize for egress fees in isolation. Optimize for the total cost of a training epoch, which includes GPU idle time. A 20% idle rate on a Lambda Labs cluster is almost certainly more expensive than the S3 egress you're trying to avoid. So before you commit to a custom cache layer, test a few object stores that explicitly advertise zero-egress and low-latency reads. Measure TTFB under sustained load, not just in a quick benchmark. And if you can't find one that works, then build the cache, but build it as a thin, disposable layer in front of your existing S3 bucket, not as a permanent part of your architecture. The goal is to get your GPUs fed, not to win a pricing argument on paper.