QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Albert Tseng, Jerry Chee, Qingyao Sun, Volodymyr Kuleshov, Christopher De Sa
Abstract
Post-training quantization (PTQ) reduces the memory footprint of LLMs by quantizing their weights to low-precision. In this work, we introduce QuIP#, a weight-only PTQ method that achieves state-of-the-art results in extreme compression regimes (≤ 4 bits per weight) using three novel techniques. First, QuIP# improves QuIP's (Chee et al., 2023) incoherence processing by using the randomized Hadamard transform, which is faster and has better theoretical properties. Second, QuIP# uses vector quantization to take advantage of the ball-shaped sub-Gaussian distribution that incoherent weights possess: specifically, we introduce a set of hardware-efficient codebooks based on the highly symmetric E 8 lattice, which achieves the optimal 8-dimension unit ball packing. Third, QuIP# uses fine-tuning to improve fidelity to the original model. Our experiments show that QuIP# outperforms existing PTQ methods, enables new behaviors in PTQ scaling, and supports fast inference. Our code can be found at https://github.com/ Cornell-RelaxML/quip-sharp .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7a155a21-2486-4cf4-a9ff-7b5e9fdff12cCited by top-tier papers138
- FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precisionJay Shah, Ganesh Bikshandi, Ying Zhang, Vijay Thakkar et al.NeurIPS 2024 · 727 citations
- QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMsSaleh Ashkboos, Amirkeivan Mohtashami, Maximilian L. Croci, Bo Li et al.NeurIPS 2024 · 723 citations
- DuQuant: Distributing Outliers via Dual Transformation Makes Stronger Quantized LLMsHaokun Lin, Haobo Xu, Yichen Wu, Jingzhi Cui et al.NeurIPS 2024 · 206 citations
- Extreme Compression of Large Language Models via Additive QuantizationVage Egiazarian, Andrei Panferov, Denis Kuznedelev, Elias Frantar et al.ICML 2024 · 187 citations
- OneBit: Towards Extremely Low-bit Large Language ModelsYuzhuang Xu, Xu Han, Zonghan Yang, Shuo Wang et al.NeurIPS 2024 · 110 citations
Builds on12
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- Up or Down? Adaptive Rounding for Post-Training QuantizationMarkus Nagel, Rana Ali Amjad, Mart van Baalen, Christos Louizos et al.ICML 2020 · 816 citations
- A Simple and Effective Pruning Approach for Large Language ModelsMingjie Sun, Zhuang Liu, Anna Bair, J. Zico KolterICLR 2024 · 794 citations
- QuIP: 2-Bit Quantization of Large Language Models With GuaranteesJerry Chee, Yaohui Cai, Volodymyr Kuleshov, Christopher De SaNeurIPS 2023 · 503 citations
Related papers
- QTIP: Quantization with Trellises and Incoherence ProcessingAlbert Tseng, Qingyao Sun, David Hou, Christopher De SaNeurIPS 2024 · 101 citations
- PV-Tuning: Beyond Straight-Through Estimation for Extreme LLM CompressionVladimir Malinovskii, Denis Mazur, Ivan Ilin, Denis Kuznedelev et al.NeurIPS 2024 · 69 citations
- Q-Palette: Fractional-Bit Quantizers Toward Optimal Bit Allocation for Efficient LLM DeploymentDeokjae Lee, Hyun Oh SongNeurIPS 2025 · 9 citations
- MLWQ: Efficient Small Language Model Deployment via Multi-Level Weight QuantizationChun Hu, Junhui He, Shangyu Wu, Yuxin He et al.EMNLP 2025 · 1 citation
- NeUQI: Near-Optimal Uniform Quantization Parameter Initialization for Low-Bit LLMsLi Lin, Xinyu Hu, Xiaojun WanICML 2026 · 1 citation
