Discover how geometric flow enables sequence models to extrapolate beyond training.

In my latest research, I introduce Geometric Flow Networks (GFN), an innovative alternative to traditional attention-based sequence modeling.

3 min readMachine Learning

This is the kind of result that challenges how we think about what a model is actually learning. A 3,164-parameter network trained on sequences of length 20 achieves perfect accuracy on sequences of length 1,000,000. That is not pattern matching. That is structure discovery. The researcher behind Geometric Flow Networks has demonstrated something worth paying attention to: an architecture that learns the underlying invariant, in this case, the toroidal symmetry of parity conservation, rather than memorizing statistical correlations. For anyone who has ever felt trapped by the scaling laws of attention-based models, this suggests a different path forward.

What makes this practical is not just the extrapolation. It is the behavior when the model fails. The Multi-Needle-in-a-Haystack test showed that with three needles, the model fires on the second needle and stops. That is a deterministic, geometrically traceable failure, not a stochastic one. You can inspect the trajectory and understand why it happened. Compare that to a transformer, where failure modes are probabilistic and opaque. For users building systems that need to be audited or debugged, this is a meaningful difference. The math behind it is not simple, but the outcome is: you get a model whose limits you can map, not just guess at.

The Inertial State Network realization on TinyShakespeare reinforces the point. With 363,000 parameters and a constant 2 KB state memory, it achieves a character-level perplexity of 2.48. That is competitive with much larger models. The caveats are honest: it was trained at length 128, so it loses coherence beyond that, and it has trouble with punctuation. But those are training scale issues, not architectural limits. The architecture itself maintains O(1) memory regardless of context length. No KV-cache, no quadratic attention, no growing state. That is a hard constraint that most sequence models cannot claim.

We think this work deserves a serious look from the research community. It is not a claim that transformers are obsolete. It is a demonstration that a physically grounded approach, second-order dynamics with symplectic integration and energy conservation, can produce behavior that statistical models cannot replicate. The code and models are public. The experiments ran on a single GTX 1650. The question now is whether others will explore the geometry enough to confirm or challenge these results. That is the only way to move forward.

From Machine Learning

I've been working on an alternative to attention-based sequence modeling that I'm calling Geometric Flow Networks (GFN). The core idea: instead of computing statistical correlations over a sequence, treat computation as a particle flowing through a geometric manifold where inputs act as perturbations that curve the trajectory without replacing the state*.* This gives three theoretical properties: O(1) state memory regardless of context length (no KV-cache), an inductive bias toward learning structural invariants rather than statistical patterns, and deterministic failure modes that are geometrically traceable rather than stochastic.

The result I can't explain away statistically:

Read the original at Machine Learning