The question posted by u/zdeneklapes lands at the exact intersection where practical training workflows meet the hard limits of distributed systems. Single-GPU auto batch finding is a convenience we have all leaned on when memory pressure builds unexpectedly. But moving that same safety net to multi-GPU setups with FSDP2 is a different beast entirely. The core issue is not about math or model architecture; it is about failure recovery semantics. When one process hits a CUDA OOM, the entire orchestration layer has to decide whether to tear down, adjust, and restart, or to fail fast and let the user intervene. Accelerate, as it stands, does not offer a first-class, supported path for that loop with FSDP2, and that gap is not a minor oversight.
We have seen this pattern before in other domains of applied machine learning. When a tool becomes powerful enough to handle real workloads, the friction points shift from raw capability to operational resilience. The ICLR Submissions Exposed: Addressing Data Privacy Concerns in AI Research story reminds us that the community often finds itself reacting to infrastructure gaps rather than designing for them from the start. Similarly, the Exploring Real-World Computer Vision: Deployments, Edge Models, and Current Challenges piece highlights how deployment constraints force engineers to build custom solutions when off-the-shelf tools stop short. Automatic batch detection across distributed processes is one of those gaps that feels like a nice-to-have until you are the person staring at a multi-hour training run that just died on the third GPU.
What is the practical path forward for someone in this exact position? You have a few honest options. You can implement a retry loop externally, where you catch the OOM at the process level, reduce the batch size globally, and relaunch the run. That works, but it is clunky and does not account for heterogeneous memory pressure across devices. Alternatively, you can manually tune the batch size using a conservative estimate from the start, accepting lower throughput for the sake of stability. The most robust approach, though, is to treat this as a known limitation and design your training pipeline with a fixed, well-tested batch size that leaves headroom for your specific model and sequence length. The Explore the Forrester Function: Beyond Mathematics, a Tool for Machine Learning piece touches on how optimization problems often reward those who understand the function they are working with rather than blindly searching for a universal solver. The same logic applies here: know your memory budget, profile it once, and stop pretending the framework will save you from yourself.
The honest take is that Accelerate and FSDP2 are not there yet for fully automatic batch recovery, and pretending otherwise will cost you more time than it saves. The community is asking the right question, but the answer is not a single flag you flip. It is a combination of disciplined profiling, external retry logic, and a willingness to accept a slightly conservative batch size. The concrete detail to watch is whether the Accelerate maintainers decide to expose a hook for distributed OOM handling in a future release. Until then, treat this as a known constraint and plan your runs accordingly. If you need this behavior yesterday, you are better off building a small wrapper around your training loop than waiting for the framework to read your mind.