Small transformers can learn to steer a flock. That is the quiet, surprising finding from a developer who built a 414,000-parameter model from scratch, trained it on 12 simulated birds, and watched it predict the movements of 50 birds it had never seen. The model hit an R² of 0.99 on held-out clips. A simpler network that ignored bird positions and attention mechanisms scored only 0.40. The difference is attention, and that difference matters far beyond a laptop simulation.
This experiment is a clean, practical demonstration of why attention-based architectures outperform older approaches at understanding relationships between entities. The model inferred the classic boid rules, cohesion, alignment, separation, without ever being told they existed. A linear probe could read cohesion from the hidden states before training even began. Alignment became readable after the first attention layer. When the researcher forced the second attention layer to attend only to itself, alignment disappeared entirely across all four runs. That is not a bug. It is evidence that the transformer's layered structure is what enables it to learn relational logic, not just memorize sequences.
The parallels to larger industry trends are hard to miss. TypeSafe's Jev hits $7.5B valuation by outpacing LLMs with fewer tokens because it achieves strong results with smaller, more efficient models. The murmuration experiment makes the same point at a microscopic scale: you do not need billions of parameters to model complex multi-agent behavior. You need the right architecture and enough data to let it discover the rules itself. And when those rules break down under scrutiny, as alignment did when the second layer was isolated, you get a window into what the model actually learned, not just what it outputs.
There is a second lesson here that is easy to overlook. The researcher's first model failed because it ignored the flock entirely. With eight ticks of history, the previous answer was already in the input, and a network that looked at only one bird outperformed the full transformer. That is a reminder that bigger models and more data do not guarantee better reasoning. A model that can cheat by repeating the past will do so unless the task forces it to attend to others. The same dynamic appears in larger AI systems, where models learn shortcuts instead of understanding. LMArena's $200M raise brings sharper focus to AI honesty and alignment is about measuring whether models are actually aligned with what users need, not just whether they produce plausible answers. The murmuration experiment shows how easy it is to build a model that looks smart but is actually just copying its own history.
The open question remains: is it fair to test a model on inputs it never saw during training, like forcing a layer to attend only to itself? The researcher admits uncertainty, and that uncertainty is worth sitting with. If the decay in alignment is caused by feeding the model a structure it never learned to process, then the test may reveal brittleness rather than a fundamental limitation. If the decay is real, it suggests attention layers are not redundant, each one builds on the last, and removing one collapses the relational understanding. Either way, the experiment provides a concrete method for probing what these models actually know. That is the detail to watch: not whether a small transformer can fly a flock, but whether we can trust what it learned when we look inside.
