Extreme Compression of Large Language Models via Additive Quantization
Vage Egiazarian, Andrei Panferov, Denis Kuznedelev, Elias Frantar, Artem Babenko, Dan Alistarh
Abstract
The emergence of accurate open large language models (LLMs) has led to a race towards performant quantization techniques which can enable their execution on end-user devices. In this paper, we revisit the problem of "extreme" LLM compression-defined as targeting extremely low bit counts, such as 2 to 3 bits per parameter-from the point of view of classic methods in Multi-Codebook Quantization (MCQ). Our algorithm, called AQLM, generalizes the classic Additive Quantization (AQ) approach for information retrieval to advance the state-of-the-art in LLM compression, via two innovations: 1) learned additive quantization of weight matrices in input-adaptive fashion, and 2) joint optimization of codebook parameters across each transformer blocks. Broadly, AQLM is the first scheme that is Pareto optimal in terms of accuracy-vs-model-size when compressing to less than 3 bits per parameter, and significantly improves upon all known schemes in the extreme compression (2bit) regime. In addition, AQLM is practical: we provide fast GPU and CPU implementations of AQLM for token generation, which enable us to match or outperform optimized FP16 implementations for speed, while executing in a much smaller memory footprint.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8542e346-6e69-4696-a979-3d00800b73c1Cited by top-tier papers91
- QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMsSaleh Ashkboos, Amirkeivan Mohtashami, Maximilian L. Croci, Bo Li et al.NeurIPS 2024 · 723 citations
- SqueezeLLM: Dense-and-Sparse QuantizationSehoon Kim, Coleman Hooper, Amir Gholami, Zhen Dong et al.ICML 2024 · 306 citations
- QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice CodebooksAlbert Tseng, Jerry Chee, Qingyao Sun, Volodymyr Kuleshov et al.ICML 2024 · 295 citations
- Exploiting LLM QuantizationKazuki Egashira, Mark Vero, Robin Staab, Jingxuan He et al.NeurIPS 2024 · 104 citations
- QTIP: Quantization with Trellises and Incoherence ProcessingAlbert Tseng, Qingyao Sun, David Hou, Christopher De SaNeurIPS 2024 · 101 citations
Builds on9
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- Up or Down? Adaptive Rounding for Post-Training QuantizationMarkus Nagel, Rana Ali Amjad, Mart van Baalen, Christos Louizos et al.ICML 2020 · 816 citations
- QuIP: 2-Bit Quantization of Large Language Models With GuaranteesJerry Chee, Yaohui Cai, Volodymyr Kuleshov, Christopher De SaNeurIPS 2023 · 503 citations
- Optimal Brain Compression: A Framework for Accurate Post-Training Quantization and PruningElias Frantar, Dan AlistarhNeurIPS 2022 · 440 citations
Related papers
- PV-Tuning: Beyond Straight-Through Estimation for Extreme LLM CompressionVladimir Malinovskii, Denis Mazur, Ivan Ilin, Denis Kuznedelev et al.NeurIPS 2024 · 69 citations
- ACBQ: Adaptive Cross-Block Quantization of Large Language ModelsHailing Wang, Jianglin Lu, Yitian Zhang, Huimin Zeng et al.ACL 2026
- LittleBit: Ultra Low-Bit Quantization via Latent FactorizationBanseok Lee, Dongkyu Kim, Youngcheon You, Youngmin KimNeurIPS 2025 · 14 citations
- LBLLM: Lightweight Binarization of Large Language Models via Three-Stage DistillationSiqing Song, Chuang Wang, Yong Lang, Yi Yang et al.ACL 2026
- LC-QAT: Data-Efficient 2-Bit QAT for LLMs via Linear-Constrained Vector QuantizationHaoyu Wang, Xingyu Yu, Haiyan Zhao, Fengxiang Wang et al.ICML 2026 · 1 citation
