linear attention

Scaling DNA modeling with linear attention that remembers what matters

A 25% recall score on a four-token DNA vocabulary isn't just a stumble; it's a sign that the compressed state in linear attention is dropping the thread.

4 min readMachine Learning

A developer working on DNA sequence modeling recently ran into a wall that many of us will recognize. They built a linear attention model to handle million-token sequences, because standard softmax attention buckles under that weight. The model did fine on standard benchmarks. Then came the needle-in-a-haystack test, and recall collapsed to around 25 percent, which is random chance for a four-token DNA alphabet. They tried tweaks, they tried existing fixes, they even ran HyenaDNA on the same test, and it also landed at 25 to 27 percent. A tiny 16K-context version managed 50 to 60 percent, but scale made the problem worse, not better.

This is not a bug report. It is a signal about the limits of compressed-state architectures. Linear attention works by folding past information into a fixed-size state, and that state has to decide what matters. For long-range recall, the state is being asked to hold a needle in a haystack that keeps growing. The developer's own experiments confirm it: modifying the architecture bought them a few points, still basically chance. The honest question they are asking is whether this is a fundamental limitation of the approach or just an engineering gap. We think the evidence leans toward fundamental, at least for the current generation of linear attention. The state is a bottleneck, and no amount of clever initialization fixes that.

What does this mean for you, if you are building with these models? It means you should not assume that a linear attention model will handle long-context retrieval just because it handles long-context inputs. The model can process a million tokens, but processing is not the same as retrieving. For DNA, where a single relevant motif might sit tens of thousands of tokens away from where it matters, this is not a corner case. It is the whole job. The developer's instinct to avoid softmax attention is sound for cost reasons, but the tradeoff is real. You are choosing between a model that can see everything and a model that can actually use what it sees.

This is where the broader conversation about adaptive systems comes in. As Mallika Rao has pointed out in her work on recommendation systems, the real complexity often lives outside the architecture itself. The same applies here: the fix may not be another attention variant, but a different way to structure the task, like external memory or retrieval-augmented steps that break the sequence into searchable pieces. And as Anthropic's work on human oversight in biology shows, sometimes the most effective approach is to keep a human in the loop rather than pushing full automation. For DNA modeling, that could mean designing benchmarks that test retrieval explicitly, not just perplexity.

The open question is whether anyone can build a linear attention variant that preserves recall at scale without sliding back toward softmax costs. That is the hard problem worth solving. The developer's 16K model worked. The 1M model did not. Somewhere between those two numbers is the threshold where the state gives up, and we do not yet know how to push it. That is the detail to watch. If someone figures that out, long-context DNA modeling changes overnight. Until then, treat linear attention for million-token sequences as a promising idea with a known failure mode, and plan your benchmarks accordingly.

From Machine Learning

Recently, I started working on DNA sequence modeling and decided to explore linear attention, mainly because DNA sequences can easily reach 1M tokens, making standard softmax attention extremely expensive in terms of memory and computation.

The model performed reasonably well on several benchmarks, but I ran into a major problem with long-range recall. On a Needle in a Haystack-style benchmark, my model was performing around 25% or even below, which is essentially random chance for a four-token DNA vocabulary (A/C/G/T).

Read the original at Machine Learning