paper web signal

HyQuant claims 3.58× decode-kernel speedup at 32K context

TL;DR

  • Top 5% of non-window key positions plus a 128-token local window cover 82.53%–85.63% of total attention mass, per the paper.
  • Decode-kernel speedup runs 1.32× to 3.58× over baseline, peaking at 32K prefix; end-to-end decode gains land at 1.04× to 1.17×.
  • The vertical-line identification step adds 3% to 5% of total runtime and lifts memory 17.4%–24.4% above a strict 4-bit design.

HyQuant, a preprint led by Jiatong Ding and colleagues, argues that most of the accuracy pain in low-bit attention comes from a small, identifiable slice of the KV cache, and that keeping just that slice in high precision buys back nearly all the quality.

The structural claim is precise. Across Qwen3-8B, Qwen3-32B, Llama-3.1-8B-Instruct and GLM-4-9B-0414, the paper reports that the top 5% of non-window key positions plus a 128-token local sliding window cover between 82.53% and 85.63% of total attention mass. Everything else is quantized to low-bit. Those regions are, the paper writes, "selected using lightweight vertical-line-aware attention-pattern signals." Identifying them adds 3% to 5% of total runtime.

Speed gains are uneven depending on where you look. The kernel benchmarks come in at "1.32× to 3.58× decode-kernel speedup," peaking at 3.58× at a 32K prefix length. End-to-end decode speedup is a more modest "1.04× to 1.17×," and the memory footprint sits 17.4% to 24.4% above a strict 4-bit design.

On accuracy the abstract stops at "nearly lossless accuracy with an extremely simple design" without publishing per-task LongBench v1 or GSM8K numbers there. Code has been released on GitHub.

Shared on Bluesky by 1 AI expert