ASRN

How language models learn to copy context with hash tables

Language models often struggle to recall information from earlier in a sequence.

3 min readMachine Learning
How language models learn to copy context with hash tables
ASRN Adaptive Sparse Recurrence Network [N]

The copy layer that learns to index context with hash tables is one of those ideas that feels obvious only after someone else has done it. And that is precisely the kind of innovation worth paying attention to. By finding earlier occurrences of the current context and copying what came next, this approach achieves memory that scales linearly with sequence length, not quadratically. For anyone who has watched a language model struggle to recall a detail from a few paragraphs back, the practical implications are immediate: longer, more coherent reasoning without the computational blowup that currently limits in-context learning.

This matters because the way we interact with data tools, from spreadsheets to AI systems, is defined by how well they remember what we have already done. The frustration of a traditional spreadsheet that cannot find the pattern you entered two sheets ago is not so different from a language model that loses the thread of a conversation after a few thousand tokens. Both break your flow. Both force you to repeat yourself. And both are problems that better memory structures can solve. The hash-table approach is elegant because it does not try to be clever about *what* to remember; it simply learns where to look for the answer it has already seen. That is a fundamentally human-centered design principle: reduce the cognitive load of repetition so users can focus on the new.

We have written before about the friction of manual workflows, such as in our piece on stop digging through folders and paste your save path directly. That same frustration applies here. When a model must reprocess context from scratch or rely on a fixed-size attention window, the user pays the price in latency and lost nuance. A copy layer with linear memory does not just improve the model; it improves the experience of using the model. It makes the tool feel more responsive to the user's actual work, which is the only benchmark that ultimately matters. Similarly, our exploration of adversarial objectives reminds us that the most productive advances in AI often come from rethinking constraints rather than adding parameters. The hash-table copy layer is exactly that kind of constraint-driven innovation.

The specific consequence to watch is how this technique interacts with retrieval-augmented generation. If a model can copy from its own context with linear memory, the line between what it "remembers" and what it "retrieves" begins to blur. That could reduce the need for external retrieval pipelines in many applications, simplifying the stack for practitioners. It also raises an open question: when a model copies from context, does it understand the copied content, or is it just efficient pattern matching? The answer will determine whether this technique is a stepping stone or a destination. Either way, it gives us a concrete reason to stop asking how much memory a model has and start asking how well it uses the memory it already holds.

From Machine Learning

A copy layer for language models that finds earlier occurrences of the current context with learned hash tables and copies what came next — with memory linear in sequence length.

Read the original at Machine Learning