Parallel-in-Time Training

Training chaotic systems in parallel: a faster path to neural network convergence

Training nonlinear RNNs on chaotic time series usually means choosing between slow sequential computation or unstable parallel methods.

3 min readMachine Learning
Training chaotic systems in parallel: a faster path to neural network convergence
Parallel-in-Time Training of Recurrent Neural Networks for Dynamical Systems Reconstruction [R]

There is a practical limit to how fast you can train a neural network on a long time series, and until recently, chaotic dynamics were that limit. The work from the NeurIPS 2026 spotlight on parallel-in-time training of recurrent neural networks for dynamical systems reconstruction does not just chip away at the problem; it demonstrates that combining two existing mechanisms can yield a speedup of more than one hundred times on sequences exceeding one million time steps. That is a concrete, measurable improvement that changes what kinds of problems become feasible to tackle.

The core insight is straightforward but the execution is anything but trivial. The DEER method already allowed the forward pass of an RNN to scale as O[(log T)²] instead of O[T] by solving fixed-point iterations in parallel across the entire sequence length. The catch was that chaotic dynamics broke the method, causing its runtime to degrade to O[T log T]. The team solved this by introducing generalized teacher forcing (GTF), which stabilizes the iterations and prevents divergence. The result is a training procedure that can handle the kind of long, unstable time series that come from real-world chaotic systems, weather data, financial signals, biological rhythms, without collapsing into exponential error growth. For anyone who has struggled to train state space models on long sequences, this feels like a direct answer to a frustration that has been treated as inevitable.

This is not an isolated advance. It sits alongside other recent work that challenges the assumption that more complexity is the only path forward. Consider the Reimagining attention with a simpler, faster Gaussian approach, which shows that a cleaner mathematical formulation can outperform the standard scaled dot-product attention mechanism. Both papers share a philosophy: look at the bottleneck, question whether the dominant approach is actually necessary, and then prove that a simpler or more targeted method works better. Similarly, the Tiered optimizer cuts MoE training memory demands by 97 percent demonstrates that memory constraints in large models can be circumvented with clever scheduling rather than brute-force hardware scaling. The pattern across these results is that the field is maturing past the era of throwing compute at problems and into a phase of architectural and algorithmic precision.

The practical consequence for anyone building with RNNs or state space models is that chaotic systems are no longer a hard wall. If you work with sensor data from physical systems, financial markets, or any domain where the underlying dynamics are sensitive to initial conditions, you now have a training method that converges reliably and quickly on sequences that were previously too long to process. The open question is how far this combination of DEER and GTF generalizes beyond the dynamical systems reconstruction setting. The paper shows huge gains against Mamba and other state space models in that specific benchmark, but the real test will be whether the same parallel-in-time stability holds when the model architecture changes or when the data has different noise characteristics. That is the detail to watch: not whether it works in the paper, but whether it works in your pipeline.

From Machine Learning

Can training of nonlinear RNNs be efficiently parallelized, ensuring fast convergence even on very long time series from chaotic systems?

In our #NeurIPS2026 spotlight “Parallel-in-Time Training of Recurrent Neural Networks for Dynamical Systems (DS) Reconstruction (DSR)” (preprint: https://arxiv.org/abs/2605.12683) we speed up training of nonlinear RNNs on time series from chaotic DS by more than 2 orders of magnitude (>100x) by combining DEER with generalized teacher forcing (GTF).

Read the original at Machine Learning