NAMOH ties attention parameter scaling to context scaling
TL;DR
- NAMOH activates K of H attention heads per token, with each head attending only to its routed subsequence instead of the full token history.
- At fixed K, raising H shortens each head's history and cuts per-token KV access without growing total KV storage.
- The authors report NAMOH outperforms fully activated models at the same total parameter count and stays compatible with GQA.
The claim: attention parameter count and usable context length can scale together without a KV cache blowup. In a paper posted to arXiv on September 30, Zizhuo Fu, Runsheng Wang and Meng Li introduce NAMOH, a sparse attention mechanism that activates K of H attention heads per token and gives each head only its routed slice of the token stream.
"Each head retains only its assigned tokens and performs causal attention within this subsequence," the authors write. At fixed K, raising H "shortens head histories and reduces per-token key-value (KV) access without increasing total KV storage." They report that NAMOH "can outperform fully activated models with the same total parameters" and runs "more efficient long-context inference than smaller dense models with matched active parameter counts," and say the design "remains compatible with GQA and existing sparse attention mechanisms."
The abstract publishes no benchmark numbers, no model scale, and no chosen values for K or H. Two researchers on our radar shared the link soon after it posted.
Shared on Bluesky by 2 AI experts
Originally reported by arxiv.org
Read the original article →Original headline: Scaling Parameter and Context in Attention: Native Sparse Attention from Mixture-of-Head