Microsoft's SWA Paper Upends Linear Attention: 60 Tokens of Context Beat Months of Post-Training
A Microsoft Applied Sciences team shows a training-free sliding-window attention mask with 4 attention sinks matches or beats post-trained linear attention across Llama, Qwen, and Phi-4 — 2-10x higher on long-context reasoning.
The linear attention bandwagon just hit a speed bump, and it arrived in the unlikeliest form: a baseline that costs nothing to deploy.
A four-person team at Microsoft’s Applied Sciences group — Alexia Jolicoeur-Martineau, Rhea Sanjay Sukthanker, Pashmina Cameron, and Emy Gervais — published a paper on arXiv titled “Sliding-window beats linear attention” (arXiv:2608.28444, submitted August 28, 2026). Its central claim is blunt enough to be a slogan: before you spend GPU-months retrofitting your LLM with linear attention, try simply masking the KV cache to a 64-token window with 4 attention sinks. The simple mask performs as well or better — and on long-context reasoning it is not even close.
The quadratic problem, reframed
Transformers pay a quadratic tax on attention. Every token generated forces the model to store its keys and values in the KV cache indefinitely, so memory and energy costs grow with context length until they become unsustainable. The industry’s favored fix over the past two years has been linear attention — a family of architectures that replaces the softmax attention matrix with constant-size state, promising linear-time inference and state-of-the-art quality “at low cost.”
The catch, the authors argue, is that this line of research “has not been properly compared to simpler baselines.” Linear attention retrofit papers typically benchmark against full attention, not against the humble sliding-window mask that anyone can apply to an existing checkpoint at inference time. That comparison gap is exactly what the paper closes.
What they tested
Sliding Window Attention (SWA) is not new — it has shipped in production models like Mistral and Gemma for years. The twist here is the combination with “attention sinks,” the technique popularized by the StreamingLLM line of work, which keeps the first few tokens of every sequence in the attention window no matter how far generation has traveled. Attention sinks stabilize the model’s output distribution, and the paper’s configuration is minimal: a window of w=64 tokens with s=4 sinks, meaning at any point the model attends only to the first 4 tokens plus the 60 most recent ones.
The configuration is entirely training-free. No fine-tuning, no distillation, no architecture surgery — you change the attention mask and the KV cache eviction policy, and you are done.
The team evaluated this against post-trained linear attention models across four instruction-tuned model families: Llama 3.1 (8B and 70B), Qwen 2.5 (32B and 72B), and Phi-4 (1.3B and 8B). The linear attention baselines included post-training methods such as LoLCATs, which converts pretrained transformers to linear attention.
The results: 99% recovery, 2-10x on long context
On short-context general knowledge benchmarks — MMLU, ARC, HellaSwag — SWA demonstrated what the authors call “remarkable robustness.” Across the board, SWA with w=64 and s=4 recovered an average of 99.0% of the original full-attention model’s performance. In 9 out of 11 evaluated scenarios, SWA beat every post-trained linear attention model it was compared against.
The specifics are stark. On MMLU, SWA recovered 93.2% of original performance — a figure the authors contrast directly against LoLCATs, the strongest linear attention retrofit, which fared worse. And SWA required zero training compute to get there.
The long-context results are where the paper earns its title. On Needle-in-a-Haystack and BABILong — the standard retrieval and multi-hop reasoning suites for long-context evaluation — SWA scored 2 to 10 times higher than post-trained linear attention. This is the counterintuitive part: a model that can only see 64 tokens at a time outperforms architectures theoretically designed to compress unbounded context into constant memory. The likely explanation is that linear attention’s compressed state loses too much information for precise retrieval, while SWA’s sinks plus recency window preserve exactly what these benchmarks demand — although the flip side, which the authors acknowledge implicitly, is that SWA genuinely cannot recall tokens outside its window, so tasks requiring true long-range recall remain out of its reach.
Why this matters
The timing is pointed. Linear attention has been enjoying a wave of enthusiasm as labs hunt for cheaper long-context inference, and several retrofit methods have been marketed as drop-in upgrades for existing checkpoints. If a free, training-free mask matches those results, the economics of that entire subfield shift. As the paper puts it: “To reduce inference memory cost, we strongly recommend switching to SWA instead of post-training linear models.”
The authors leave a door open for linear attention — the closing line concedes that linear models “likely require to be trained from scratch or extensive post-training in order to even match SWA.” That is a defensible scientific position and simultaneously a serious indictment: if your cheap retrofit loses to a 64-token window, the retrofit is not ready.
Jolicoeur-Martineau, known for the Tiny Recursive Models work, announced the paper on X with the pithiest version of the claim: “switching to a sliding-window attention mask with attention sinks (at no cost) beats linear attention post-training.” The paper has been climbing aggregator leaderboards since, sitting in the top ten on cs.CL daily listings and drawing heavy discussion on r/LocalLLaMA, where practitioners immediately clocked the practical implications for local inference setups.
There are honest caveats. A 64-token window is an aggressive setting chosen to stress-test the baseline; real deployments may prefer larger windows that trade memory for recall. SWA cannot do true arbitrary-position retrieval, so anyone needing faithful 1M-token context will not find it here. And the linear attention field is moving fast — methods trained from scratch, rather than retrofitted, remain genuinely competitive.
But as a corrective to benchmark hygiene, the paper is a small classic: before you architect your way out of a problem, check whether a mask solves it. For an industry currently pouring billions into long-context infrastructure, “the boring baseline still wins” is a message worth hearing — especially when the baseline is free.