RNNs

Mapping AI's Memory: Where RNNs, Transformers, and SSMs Store What They Know

If you've ever wondered why some AI models seem to remember everything while others compress the past into a tight hidden state, the answer is all about where memory lives.

3 min readMachine Learning

The most productive question in AI architecture right now isn't "which model wins?" but "where does the memory actually live?" For anyone building on top of foundation models, whether you're deploying a chatbot, fine-tuning a spreadsheet copilot, or experimenting with long-context data pipelines, this distinction matters more than another benchmark score. The memory bottleneck is the practical limit on how far your work can scale.

A thoughtful analysis of RNNs, Transformers, and state-space models (SSMs) lays out the trade-offs with unusual clarity. Recurrent networks compress everything into a single hidden state, elegant but fundamentally constrained: a model with O(N²) parameters can only carry O(N) bits of state forward. Transformers sidestep this by storing past tokens as key-value entries, little post-it notes that accumulate with every step. That gives them extraordinary recall, but at the cost of a growing cache that demands more memory and compute as context lengthens. As one recent exploration noted, Five startups earning trust and traction at PearX demo day are already building products that depend on managing this exact tension between context and efficiency.

What makes this conversation timely is the emergence of selective SSMs like Mamba and hybrid architectures such as Dragon Hatchling (BDH). These systems bring back fixed-size recurrent memory but with input-dependent retention: the model decides what to keep or forget based on each incoming token. Significantly, BDH stores its recurrent state as an N×D matrix in high-dimensional neuron space rather than materializing the full N×N connectivity matrix. That moves working memory closer to the model's internal synaptic structure. Is this a fundamental advance or just a more clever compression scheme? The answer has direct consequences. Projects that need durable, long-term understanding, like TypeSafe AI's Jev Delivers Focused Utility Without Hallucinations, must grapple with whether memory lives in transient context or in the model's frozen weights.

The real takeaway is that every architecture faces the same underlying constraint: you either compress history into a finite state or you expand memory to hold it all. Transformers excel at recall at the cost of growth. RNNs and SSMs trade perfect recall for bounded, efficient state. There is no free lunch. But reframing the question from "which architecture wins?" to "where does memory live?" pushes us past the horse race and toward a more pragmatic evaluation. Watch for architectures that blur the line between working memory and learned connectivity, that is where the next practical leap will come from, not from another headline about parameter counts.

From Machine Learning

Someone who has always loved looking at the space between different AI techniques, this time I went a little deeper into the memory trade-offs between RNNs, Transformers and SSMs. I found it interesting because once you start looking at these architectures through the lens of working memory, a lot of the differences become easier to understand. Where does the memory actually live? Is it a compact recurrent state, a growing KV cache, or something closer to the network itself?

Read the original at Machine Learning