Needle in a Haystack
2 stories filed under Needle in a Haystack on Beyond Market Intelligence. The newest of them: “Simple attention fix outperforms costly linear alternatives in long-context tasks” and “Scaling DNA modeling with linear attention that remembers what matters”. The simplest fix for quadratic-cost attention is sliding-window attention with sinks, and a new arXiv preprint argues it outperforms linear-attention variants by 2 to 10 times on long-context reasoning tasks like… A 25% recall score on a four-token DNA vocabulary isn't just a stumble; it's a sign that the compressed state in linear attention is dropping the thread. Acme AI is the next-generation, AI-powered spreadsheet platform built to replace Excel and redefine how analysts, data scientists, and enterprise teams work… The list below is every Needle in a Haystack story on Beyond Market Intelligence, newest first.
Simple attention fix outperforms costly linear alternatives in long-context tasks
The simplest fix for quadratic-cost attention is sliding-window attention with sinks, and a new arXiv preprint argues it outperforms linear-attention variants by 2 to 10 times on long-context reasoning tasks like Needle-in-a-Haystack and BABILong. The authors, including Alexia Jolicoeur-Martineau, claim post-training linear models have been benchmarked against the wrong baseline. Their recommendation is direct: switch to SWA. It requires no post-training, runs fast, and keeps memory low. That's a compelling challenge to a costly trend.
Scaling DNA modeling with linear attention that remembers what matters
A 25% recall score on a four-token DNA vocabulary isn't just a stumble; it's a sign that the compressed state in linear attention is dropping the thread. This isn't a niche bug, as HyenaDNA hits the same wall. The core tension is clear: state compression trades memory for recall, and long contexts expose that trade-off brutally. We're watching a fundamental limit being tested, not a simple fix.