CBQ: Cross-Block Quantization for Large Language Models
Xin Ding, Xiaoyu Liu, Zhijun Tu, Yun Zhang, Wei Li, Jie Hu, Hanting Chen, Yehui Tang, Zhiwei Xiong, Baoqun Yin, Yunhe Wang
Abstract
Post-training quantization (PTQ) has played a pivotal role in compressing large language models (LLMs) at ultra-low costs. Although current PTQ methods have achieved promising results by addressing outliers and employing layer-or blockwise loss optimization techniques, they still suffer from significant performance degradation at ultra-low bits precision. To dissect this issue, we conducted an indepth analysis of quantization errors specific to LLMs and surprisingly discovered that, unlike traditional sources of quantization errors, the growing number of model parameters, combined with the reduction in quantization bits, intensifies inter-layer and intra-layer dependencies, which severely impact quantization accuracy. This finding highlights a critical challenge in quantizing LLMs. To address this, we propose CBQ, a cross-block reconstruction-based PTQ method for LLMs. CBQ leverages a cross-block dependency to establish long-range dependencies across multiple blocks and integrates an adaptive LoRA-Rounding technique to manage intra-layer dependencies. To further enhance performance, CBQ incorporates a coarse-to-fine pre-processing mechanism for processing weights and activations. Extensive experiments show that CBQ achieves superior low-bit quantization (W4A4, W4A8, W2A16) and outperforms existing state-of-the-art methods across various LLMs and datasets. Notably, CBQ only takes 4.3 hours to quantize a weightonly quantization of a 4-bit LLAMA1-65B model, achieving a commendable trade off between performance and efficiency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 03713b04-1995-410b-9db8-4597cbc63670Cited by top-tier papers14
- Q-VLM: Post-training Quantization for Large Vision-Language ModelsChangyuan Wang, Ziwei Wang, Xiuwei Xu, Yansong Tang et al.NeurIPS 2024 · 57 citations
- Model-Preserving Adaptive RoundingAlbert Tseng, Zhaofeng Sun, Chris De SaICML 2026 · 17 citations
- SliderQuant: Accurate Post-Training Quantization for LLMsShigeng Wang, Chao Li, Yangyuxuan Kang, Jiawei Fan et al.ICLR 2026 · 6 citations
- Learning Grouped Lattice Vector Quantizers for Low-Bit LLM CompressionXi Zhang, Xiaolin Wu, Jiamang Wang, Weisi LinNeurIPS 2025 · 4 citations
- Task-Specific Zero-Shot Quantization-Aware Training for Object DetectionChanghao Li, Xinrui Chen, Ji Wang, Kang Zhao et al.ICCV 2025 · 2 citations
Builds on21
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language ModelsGuangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu et al.ICML 2023 · 1,493 citations
- Aligning AI With Shared Human ValuesDan Hendrycks, Collin Burns, Steven Basart, Andrew Critch et al.ICLR 2021 · 878 citations
Related papers
- ACBQ: Adaptive Cross-Block Quantization of Large Language ModelsHailing Wang, Jianglin Lu, Yitian Zhang, Huimin Zeng et al.ACL 2026
- SERQ: Saliency-Aware Low-Rank Error Reconstruction for LLM QuantizationYeonsik Park, Hyeonseong Kim, Seungkyu ChoiICLR 2026 · 1 citation
- Layer-Wise High-Impact Parameter Ratio Optimization in Post-Training Quantization for Large Language ModelsCuong Pham, Dung Anh Hoang, Cuong C. Nguyen, Trung Le et al.ACL 2026
- PTQ1.61: Push the Real Limit of Extremely Low-Bit Post-Training Quantization Methods for Large Language ModelsJiaqi Zhao, Miao Zhang, Ming Wang, Yuzhang Shang et al.ACL 2025 · 7 citations
- QLLM: Accurate and Efficient Low-Bitwidth Quantization for Large Language ModelsJing Liu, Ruihao Gong, Xiuying Wei, Zhiwei Dong et al.ICLR 2024 · 75 citations
