KronQ Makes 2-Bit LLaMA-3-70B Work Where GPTQ Degenerates
Summary
GPTQ and GPTAQ both collapse at 2-bit quantization of LLaMA-3-70B (perplexity >2,000); KronQ achieves 7.93—a result that makes extreme LLM compression practical on consumer hardware and materially changes inference deployment cost calculus.
Originally reported by paper
Read the original article →Original headline: KronQ Makes 2-Bit LLaMA-3-70B Work Where GPTQ Degenerates