Frozen memory, lasting recall: AI retains facts through a cold reboot.

Introducing a groundbreaking frozen transformer model with an isolated memory buffer, capable of retaining knowledge even after a cold reload.

3 min readMachine Learning

The persistence of memory in AI has always been the fragile part of the story. Models forget, drift, or collapse into the noise of their training data the moment you look away. So when a transformer freezes its backbone, isolates a memory buffer, and still answers "wombats produce cube-shaped" with "droppings" at 0.9997 probability after a cold reboot, that is not a small trick. It is a quiet proof that durable recall can be engineered without touching the weights that do the thinking. The authors of this work have done something more useful than adding another layer of capacity: they have separated the act of remembering from the act of reasoning.

What stands out is the mechanism itself. The model computes a co-activation outer product at every token step, but instead of discarding it, the system accumulates it into a learned content-addressing projection. That means the address for each memory reflects the full causal context of the sequence, not just the token in front of it. This is not a clever hack bolted onto the side of a transformer. It is a deliberate architectural choice that sits firmly in the fast weight programmer tradition, and it lands close to concurrent work from FwPKM and In-Place TTT, both of which converge on a similar write rule independently. That convergence is a signal. When three separate groups arrive at the same shape of solution, you are no longer looking at a novelty. You are looking at a pattern.

For the practitioner, the implications are immediate and practical. The system encodes 20 unrelated facts jointly with perfect accuracy, median probability at 0.997, and cross-contamination below 0.03 when two subjects are stored at once. It does this with 15 million parameters, 250 million tokens, and a single consumer GPU. The entire encoding process takes 300 gradient steps on the memory weights only, not a full fine-tune, not a retraining run. That means you can add new knowledge to a model without touching its backbone, without catastrophic forgetting, and without needing a cluster. You can run it yourself today. The code is in the README, under an Apache 2.0 license, and the instructions are direct.

The honest limitation is capacity. Twenty facts is not a knowledge base, and the authors do not claim otherwise. But the point is not the number. The point is that the mechanism works, survives a cold reload, and does not corrupt the frozen core. That is the foundation of something worth exploring. We are not saying this is the end of the road for memory in neural networks. We are saying the road just got a lot more specific. If you have been waiting for a reason to experiment with isolated memory buffers, this is it. Go run the code. See what breaks. That is how progress gets made.

From Machine Learning

A transformer with a separate, isolated memory buffer. Backbone frozen. 300 gradient steps on the memory weights only:

Save, kill process, cold reload, query again. Same result. 20 unrelated facts encoded jointly: 20/20 correct, median p = 0.997. Two subjects encoded simultaneously with cross-contamination < 0.03.

Read the original at Machine Learning