CoWindow Attention

Two new attention models slash redundancy to stretch context further

Attention research has a habit of making the same trade-off: sacrifice coverage for speed.

4 min readMachine Learning
From Machine Learning

I'm one of the authors of two recent papers exploring different sources of redundant computation in attention. I'd like to share the ideas and hear feedback from people working on long-context models and attention kernels.

CoWindow Attention (CoWA) distributes distant context across KV heads using complementary windows, while sharing local and prefix-sink windows. Each head attends sparsely, but the union of their visible positions covers the full causal history. The pattern is position-defined and requires no learned router or indexer.

Read the original at Machine Learning