KLQ Ships Training-Free Rotation Quantization That Beats Prior W4A4KV4 Methods on Llama 3.2 and Qwen 2.5
Summary
KLQ, released on GitHub and shared to r/LocalLLaMA, eigendecomposes activation covariance to pinpoint sensitive directions and water-fills bit budget across them — no calibration training required. At W4A4KV4, KLQ-RTN cuts Qwen 2.5 0.5B perplexity to 21.07 vs CoQuant's 219.9, and its VQ variant matches or beats ReSpinQuant on Llama 3.2 1B. Author claims it is the first training-free method to hold its own against training-based rotation quantizers at aggressive precisions.
Originally reported by github.com
Read the original article →Original headline: KLQ Ships Training-Free Rotation Quantization That Beats Prior W4A4KV4 Methods on Llama 3.2 and Qwen 2.5