LoRA

Gradient accumulation speed varies more than expected across GPU setups

Conventional wisdom says a batch of four is a batch of four, but this test on LoRA with Qwen3-1.7B shows otherwise. The user found that on a T4, running four micro-batches before one optimizer step was 17% slower than a…

4 min readMachine Learning

There's a quiet assumption buried in a lot of machine learning work that goes something like this: if the math says the optimizer sees the same batch, the hardware should behave. That assumption is wrong, and a recent experiment posted by a user on r/MachineLearning makes the point beautifully. They ran Qwen3-1.7B with LoRA for 100 optimizer updates on both a T4 and an L4, testing three gradient accumulation schedules: 1×4, 2×2, and 4×1. The effective batch was four every time. The training time was not even close. On the T4, 4×1 was about 17% faster than 1×4. On the L4, the gap stretched to 41%. Same examples, same optimizer, same step count. The only difference was how the GPU physically received the work.

This is the kind of finding that should make you question every benchmark you have ever trusted. We often treat effective batch size as the single lever that matters for optimization behavior, and physical batch as a mere memory constraint. The experiment shows they are actually two separate choices that pull on different parts of the system. Effective batch is about the optimization trajectory. Physical batch and accumulation control execution shape, kernel tiling, launch overhead, and how well the GPU pipeline fills. The experiment puts it plainly: they are not the same knob. And the data backs that up. On the L4, 2×2 was actually slightly faster than 4×1, which means performance is not even linear as physical batch grows. The GPU is not a simple calculator. It rewards some shapes more than others, and those rewards shift across hardware generations.

This connects to a larger pattern we have been tracking in the industry. We recently covered how Beyond the Hype: Why AI "Escapes" Are Really Firewall Shortcomings shows that most AI safety narratives fail because people assume the system behaves the way the marketing deck describes it. This is the same instinct, applied to performance tuning. You cannot assume the abstraction layer tells you the whole story. The other piece, [Sharing my ML learning repo, NumPy to Transformers, 5 months, daily commits, all notebooks public. [D]](/post/sharing-my-ml-learning-repo-numpy-to-transformers-5-months-d-cmu91wx2d04i95ngm1211p2j9), is a reminder that real understanding in this field comes from grinding through the details, not from reading the docs. The Hugging Face documentation already says gradient accumulation does not improve throughput over a true larger batch. This experiment gives that warning teeth, and it gives you a practical way to act on it.

So what do we do with this? First, stop treating gradient accumulation as a free pass. It is a memory workaround, not a performance feature. If you are training on a single GPU, start with the largest physical batch that fits comfortably, then test a few nearby accumulation values. The author used TraceML to get step-level timing, and the notebook is public. That is the kind of reproducibility we need more of. The deeper takeaway is that your GPU has a personality, and you have to learn it. The difference between 238 seconds and 287 seconds on the same model is not noise. It is the difference between a training run that fits into a lunch break and one that pushes into the afternoon. For practitioners running dozens of experiments a week, that is not a micro-optimization. It is the difference between shipping and waiting. The open question is how this scales on multi-GPU setups or with different model families, but the lesson is already clear: measure the physical batch, not just the effective one. Your training time is not a given. It is a choice you are making every time you pick a number.

From Machine Learning

I had assumed 1 × 4, 2 × 2 and 4 × 1 will take somewhat similar time because effective batch is 4 in all cases.

I ran Qwen3-1.7B with TRL and LoRA for 100 optimizer updates.

Read the original at Machine Learning