paper web signal

STEPQuant matches FP32 at 6-bit, cuts serving memory by 68.7%

TL;DR

  • STEPQuant is a post-training quantization scheme that compresses Delta-rule recurrent states while matching FP32-state accuracy at a 6-bit budget.
  • On Qwen3.8-27B and Kimi-Linear-48B-A3B-Instruct, 6-bit STEPQuant delivers over 5x recurrent-state compression and up to 68.7% total serving-memory reduction in SGLang.
  • At 4-bit, the method outperforms uniform INT8 by allocating bits by error magnitude and memory lifetime across key rows and value columns.

Linear attention was supposed to kill the KV cache. The fixed-size recurrent states that replace it, the authors argue, 'can become a substantial memory bottleneck under concurrent serving.'

That observation anchors a new arXiv preprint by Bingchen Yao and colleagues introducing STEPQuant, a post-training quantization framework for Delta-rule recurrent states. The paper's core problem is error propagation: 'Directly quantizing recurrent states to low precision often leads to severe accuracy degradation, as quantization errors propagate through successive state updates.'

STEPQuant's answer is to allocate bits unevenly. The method 'allocates precision according to error magnitude and memory lifetime, and jointly fits key-row and value-column scales based on state distributions and key-row impact on output error.' Errors in long-lived memory and errors in high-impact key rows get more bits; the rest get fewer.

Tested on Qwen3.8-27B and Kimi-Linear-48B-A3B-Instruct, STEPQuant 'closely matches FP32-state accuracy under a nominal 6-bit budget' and 'outperforms uniform INT8 in its 4-bit configuration,' the paper reports. Integrated into SGLang with optimized GPU kernels, the 6-bit version delivers 'over 5x recurrent-state compression' and reduces total serving memory 'by up to 68.7%.'

The abstract publishes no per-benchmark accuracy numbers and does not quantify any throughput gain. Code is on GitHub.