BiLLM: Pushing the Limit of Post-Training Quantization for LLMs
Wei Huang, Yangdong Liu, Haotong Qin, Ying Li, Shiming Zhang, Xianglong Liu, Michele Magno, Xiaojuan Qi
Abstract
Pretrained large language models (LLMs) exhibit exceptional general language processing capabilities but come with significant demands on memory and computational resources. As a powerful compression technology, binarization can extremely reduce model weights to a mere 1 bit, lowering the expensive computation and memory requirements. However, existing quantization techniques fall short of maintaining LLM performance under ultra-low bit-widths. In response to this challenge, we present BiLLM, a groundbreaking 1-bit post-training quantization scheme tailored for pretrained LLMs. Based on the weight distribution of LLMs, BiLLM first identifies and structurally selects salient weights, and minimizes the compression loss through an effective binary residual approximation strategy. Moreover, considering the bell-shaped distribution of the non-salient weights, we propose an optimal splitting search to group and binarize them accurately. BiLLM achieving for the first time high-accuracy inference (e.g. 8.41 perplexity on LLaMA2-70B) with only 1.08-bit weights across various LLMs families and evaluation metrics, outperforms SOTA quantization methods of LLM by significant margins. Moreover, BiLLM enables the binarization process of the LLM with 7 billion weights within 0.5 hours on a single GPU, demonstrating satisfactory time efficiency. Our code is available at https://github.com/Aaronhuang-778/BiLLM.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0b44c59f-3b38-466e-83fd-6c95486b3dacCited by top-tier papers37
- PV-Tuning: Beyond Straight-Through Estimation for Extreme LLM CompressionVladimir Malinovskii, Denis Mazur, Ivan Ilin, Denis Kuznedelev et al.NeurIPS 2024 · 69 citations
- ParetoQ: Improving Scaling Laws in Extremely Low-bit LLM QuantizationZechun Liu, Changsheng Zhao, Hanxian Huang, Sijia Chen et al.NeurIPS 2025 · 50 citations
- Q-SNNs: Quantized Spiking Neural NetworksWenjie Wei, Yu Liang, Ammar Belatreche, Yichen Xiao et al.ACM MM 2024 · 23 citations
- Flow Caching for Autoregressive Video GenerationYuexiao Ma, Xuzhe Zheng, Jing Xu, Xiwei Xu et al.ICLR 2026 · 20 citations
- Mixture of Scales: Memory-Efficient Token-Adaptive Binarization for Large Language ModelsDongwon Jo, Taesu Kim, Yulhwa Kim, Jae-Joon KimNeurIPS 2024 · 14 citations
Builds on6
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- Multitask Prompted Training Enables Zero-Shot Task GeneralizationVictor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach et al.ICLR 2022 · 1,976 citations
- QuIP: 2-Bit Quantization of Large Language Models With GuaranteesJerry Chee, Yaohui Cai, Volodymyr Kuleshov, Christopher De SaNeurIPS 2023 · 503 citations
- PB-LLM: Partially Binarized Large Language ModelsZhihang Yuan, Yuzhang Shang, Zhen DongICLR 2024 · 91 citations
Related papers
- ARB-LLM: Alternating Refined Binarizations for Large Language ModelsZhiteng Li, Xianglong Yan, Tianao Zhang, Haotong Qin et al.ICLR 2025
- HBLLM: Wavelet-Enhanced High-Fidelity 1-Bit Quantization for LLMsNingning Chen, Weicai Ye, Ying JiangNeurIPS 2025 · 5 citations
- SliM-LLM: Salience-Driven Mixed-Precision Quantization for Large Language ModelsWei Huang, Haotong Qin, Yangdong Liu, Yawei Li et al.ICML 2025
- PTQ1.61: Push the Real Limit of Extremely Low-Bit Post-Training Quantization Methods for Large Language ModelsJiaqi Zhao, Miao Zhang, Ming Wang, Yuzhang Shang et al.ACL 2025 · 7 citations
- NanoQuant: Efficient Sub-1-Bit Quantization of Large Language ModelsHyochan Chong, Dongkyu Kim, Changdong Kim, Minseop ChoiICML 2026
