HyQuant claims 3.58× decode-kernel speedup at 32K context
TL;DR
- Top 5% of non-window key positions plus a 128-token local window cover 82.53%–85.63% of total attention mass, per the paper.
- Decode-kernel speedup runs 1.32× to 3.58× over baseline, peaking at 32K prefix; end-to-end decode gains land at 1.04× to 1.17×.
- The vertical-line identification step adds 3% to 5% of total runtime and lifts memory 17.4%–24.4% above a strict 4-bit design.
HyQuant, a preprint led by Jiatong Ding and colleagues, argues that most of the accuracy pain in low-bit attention comes from a small, identifiable slice of the KV cache, and that keeping just that slice in high precision buys back nearly all the quality.
The structural claim is precise. Across Qwen3-8B, Qwen3-32B, Llama-3.1-8B-Instruct and GLM-4-9B-0414, the paper reports that the top 5% of non-window key positions plus a 128-token local sliding window cover between 82.53% and 85.63% of total attention mass. Everything else is quantized to low-bit. Those regions are, the paper writes, "selected using lightweight vertical-line-aware attention-pattern signals." Identifying them adds 3% to 5% of total runtime.
Speed gains are uneven depending on where you look. The kernel benchmarks come in at "1.32× to 3.58× decode-kernel speedup," peaking at 3.58× at a 32K prefix length. End-to-end decode speedup is a more modest "1.04× to 1.17×," and the memory footprint sits 17.4% to 24.4% above a strict 4-bit design.
On accuracy the abstract stops at "nearly lossless accuracy with an extremely simple design" without publishing per-task LongBench v1 or GSM8K numbers there. Code has been released on GitHub.
Shared on Bluesky by 1 AI expert
Originally reported by paper
Read the original article →Original headline: HyQuant Reaches 3.58× FlashAttention-2 Decode Speed at 32K via Hybrid Attention Precision — EMNLP 2026