Sliding window beats linear attention
Preprint shows Sliding Window Attention outperforms linear attention by 2-10x on long-context tasks without retraining.
ImportanceMaterialEvidenceE2 unreplicatedWrite-upQuick
Sliding Window Attention (SWA) with attention sinks performs as well or better than most retrofitted Linear Attention models across multiple LLMs.
The study compares two methods for reducing LLM memory consumption: compressing the KV cache (via SWA) and retrofitting models to use Linear Attention with fixed-size states. On long-context reasoning benchmarks like Needle-in-a-Haystack and BABILong, SWA achieved 2 to 10 times higher performance than linear attention.
SWA requires no additional training, is extremely fast, and uses little memory. The authors argue that when training budgets are limited, switching to SWA is a much more effective way to reduce inference costs than retrofitting linear attention.
This is an arXiv preprint (v2 updated Oct 4, 2026); results have not yet been independently reproduced.