PyTorch

Why Your AI Model Slows 170x on a T4 vs A100

A 170x slowdown between a T4 and an A100 is not a generational gap; it's a signal that something structural is off.

4 min readMachine Learning

A 170x performance gap between an NVIDIA T4 and an A100 is not a hardware difference; it is a signal. When a point-tracking model runs at 0.5 seconds per half-video on an A100 and 85 seconds on a T4, with GPU utilization pinned at 99% on both, you are not looking at raw compute limits. You are looking at a fundamental mismatch between the model's execution pattern and the T4's architecture. The fact that pure FP32 is involved, combined with dense 4D correlation volumes and transformer layers, points to a specific culprit: memory bandwidth and the cost of non-tensor-core operations. The A100 does not just have more compute; it has a vastly wider memory bus and faster HBM2e. But even that does not explain 170x. What does is the likelihood that your kernel is memory-bound in a way that punishes the T4's smaller L2 cache and lower memory bandwidth per byte of data moved.

We have seen this pattern before in how models behave across different accelerators. The transformer layer behavior often comes down to token-wise memory access patterns, and the same principle applies here. Your 4D correlation volume is effectively a dense matching table between all pairs of frames. At 47 frames and 256x256 resolution, that volume is enormous, and building it in FP32 means writing and reading hundreds of megabytes of intermediate data. On the A100, that fits more comfortably in the 40MB L2 cache. On the T4, with its 4MB L2, you are spilling to VRAM repeatedly. That is not a generational gap; that is a cache thrashing problem. The T4 is not slow because it is old. It is slow because your model's working set does not fit in its on-chip memory, and every access to that correlation volume becomes a trip to DRAM.

What should you do first? Profile the memory traffic, not the kernel time. Use a profiler that shows memory throughput and L2 hit rates, not just GPU utilization. We would bet real money that you see a massive difference in L2 hit rate between the two cards. The fix is not to buy an A100. The fix is to reduce the precision of the correlation volume to FP16 or even BF16 for the matching computation, or to tile the 4D volume so that you process a few frame pairs at a time, keeping the working set inside the T4's cache. This is exactly the kind of optimization that the DSP-inspired semantic vocoder approach suggests: treat the model's data flow as a signal processing problem, and you can often restructure the computation to be cache-resident. The T4 is not a toy; it is a capable inference card, but it rewards models that respect its memory hierarchy.

The deeper lesson here is that hardware benchmarks are only useful when the workload matches the architecture. A 170x gap is not a reason to abandon the T4; it is a reason to question your assumptions about where the bottleneck actually lives. If we were advising you, we would say this: do not chase the A100. Chase the memory footprint. Run your model with `torch.jit.trace` and inspect the memory access graph. You will likely find that a simple change, such as fusing the correlation volume construction with the transformer attention, collapses that 170x to something closer to 10x, which is the real hardware difference. And when you find that, you will have learned more about your model than any benchmark could teach you. The next time you see a number that seems absurd, treat it as a clue, not a verdict.

From Machine Learning

Seeing a ~170× slowdown running a point-tracking model on an NVIDIA T4 compared to an A100. On A100 the tracker takes ~0.5 seconds per half-video. On T4 the same call takes ~85 seconds. Video is 47 frames at 256×256, batch 1. I expect a meaningful gap between these cards, but 170× feels too large to explain by generational hardware differences alone.

Given the architecture (4D correlations + transformers) and pure FP32 execution, what would cause a T4 to be this much slower than A100? What should I look for or profile first?

Read the original at Machine Learning