Fudan Paper Hits 16x Length Extrapolation on Hybrid LLMs
TL;DR
- A 776M hybrid trained at 4k holds 100% NIAH-SK1 accuracy out to 64k, a 16x training-free length extrapolation.
- The authors name a 'Seesaw Effect': sliding-window hybrids extrapolate farther, linear-attention hybrids benefit more from long-context continual pretraining.
- Models from 376M to 3B were tested; validation is limited to the pretraining stage, with post-training and multimodal left for future work.
The Fudan paper reports 16x training-free length extrapolation on hybrid long-context LLMs, from a 4k pretraining window out to 64k, while keeping 100% accuracy on the NIAH-SK1 retrieval test. The method, Sliding-Window Linear Attention (SWLA), was tested on models from 376M to 3B parameters.
The authors state the design prescription directly: 'position-biased attention should use a sliding window to restrict its relative-position information, while NoPE attention should strengthen its global perception.' That split is the core of their hybrid recipe.
The paper also diagnoses what it calls a 'Seesaw Effect': sliding-window attention hybrids extrapolate farther than linear-attention hybrids, while linear-attention hybrids pick up more from long-context continual pretraining. The authors describe the same tension as a 'No-Free-Lunch Effect' between the two attention families.
Short-context pretraining ran at 4k on 50B tokens; a long-context continual pretraining phase ran at 32k on 5B tokens. The headline retrieval result is for the 776M configuration labeled GLA-NoPE-LH, combined with the authors' Extrapolation based on Matthew Effect (EME) strategy. The paper's scope limits are stated plainly: validation is on the pretraining stage only, and sparse-attention and compressed-attention hybrids are left for future work. The preprint is on Hugging Face, and it lands in a busy week for long-context architecture work on our open-source tracker.
Originally reported by huggingface.co
Read the original article →Original headline: Fudan SWLA Paper Hits 16x Training-Free Length Extrapolation on Hybrid Long-Context LLMs