NestQuant: nested lattice quantization for matrix products and LLMs
Semyon Savkin, Eitan Porat, Or Ordentlich, Yury Polyanskiy
Abstract
Post-training quantization (PTQ) has emerged as a critical technique for efficient deployment of large language models (LLMs). This work proposes NESTQUANT, a novel PTQ scheme for weights and activations that is based on self-similar nested lattices. Recent works have mathematically shown such quantizers to be information-theoretically optimal for low-precision matrix multiplication. We implement a practical lowcomplexity version of NestQuant based on Gosset lattice, making it a drop-in quantizer for any matrix multiplication step (e.g., in self-attention, MLP etc). For example, NestQuant quantizes weights, KV-cache, and activations of Llama-3-8B to 4 bits, achieving perplexity of 6.6 on wikitext2. This represents more than 55% reduction in perplexity gap with respect to unquantized model (perplexity of 6.14) compared to state-of-the-art Meta's SpinQuant (perplexity 7.3), OstQuant (7.3) and QuaRot (8.2). Comparisons on bigger models (up to 70B) and on various LLM evaluation benchmarks confirm uniform superiority of NestQuant.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 92b63cbe-ccc6-4c1d-a011-7ed26ebf3cd4Cited by top-tier papers5
- Model-Preserving Adaptive RoundingAlbert Tseng, Zhaofeng Sun, Chris De SaICML 2026 · 17 citations
- WaterSIC: information-theoretically (near) optimal linear layer quantizationEgor Lifar, Semyon Savkin, Or Ordentlich, Yury PolyanskiyICML 2026 · 5 citations
- Learning Grouped Lattice Vector Quantizers for Low-Bit LLM CompressionXi Zhang, Xiaolin Wu, Jiamang Wang, Weisi LinNeurIPS 2025 · 4 citations
- UniSVQ: 2-bit Unified Scalar-Vector QuantizationHaoyu Wang, Haiyan Zhao, Xingyu Yu, Zhangyang Yao et al.ICML 2026 · 2 citations
- LC-QAT: Data-Efficient 2-Bit QAT for LLMs via Linear-Constrained Vector QuantizationHaoyu Wang, Xingyu Yu, Haiyan Zhao, Fengxiang Wang et al.ICML 2026 · 1 citation
Builds on12
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language ModelsGuangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu et al.ICML 2023 · 1,493 citations
- QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMsSaleh Ashkboos, Amirkeivan Mohtashami, Maximilian L. Croci, Bo Li et al.NeurIPS 2024 · 723 citations
- QuIP: 2-Bit Quantization of Large Language Models With GuaranteesJerry Chee, Yaohui Cai, Volodymyr Kuleshov, Christopher De SaNeurIPS 2023 · 503 citations
Related papers
- SpinQuant: LLM Quantization with Learned RotationsZechun Liu, Changsheng Zhao, Igor Fedorov, Bilge Soran et al.ICLR 2025
- ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank ResidualsUtkarsh Saxena, Sayeh Sharify, Kaushik Roy, Xin WangICML 2025
- NanoQuant: Efficient Sub-1-Bit Quantization of Large Language ModelsHyochan Chong, Dongkyu Kim, Changdong Kim, Minseop ChoiICML 2026
- Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank CompensationZhewei Yao, Xiaoxia Wu, Cheng Li, Stephen Youn et al.AAAI 2024 · 50 citations
- KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache QuantizationColeman Hooper, Sehoon Kim, Hiva Mohammadzadeh, Michael W. Mahoney et al.NeurIPS 2024 · 738 citations
