CBQ: Cross-Block Quantization for Large Language Models
Xin Ding, Xiaoyu Liu, Zhijun Tu, Yun Zhang, Wei Li, Jie Hu, Hanting Chen, Yehui Tang, Zhiwei Xiong, Baoqun Yin, Yunhe Wang
摘要
Post-training quantization (PTQ) has played a pivotal role in compressing large language models (LLMs) at ultra-low costs. Although current PTQ methods have achieved promising results by addressing outliers and employing layer-or blockwise loss optimization techniques, they still suffer from significant performance degradation at ultra-low bits precision. To dissect this issue, we conducted an indepth analysis of quantization errors specific to LLMs and surprisingly discovered that, unlike traditional sources of quantization errors, the growing number of model parameters, combined with the reduction in quantization bits, intensifies inter-layer and intra-layer dependencies, which severely impact quantization accuracy. This finding highlights a critical challenge in quantizing LLMs. To address this, we propose CBQ, a cross-block reconstruction-based PTQ method for LLMs. CBQ leverages a cross-block dependency to establish long-range dependencies across multiple blocks and integrates an adaptive LoRA-Rounding technique to manage intra-layer dependencies. To further enhance performance, CBQ incorporates a coarse-to-fine pre-processing mechanism for processing weights and activations. Extensive experiments show that CBQ achieves superior low-bit quantization (W4A4, W4A8, W2A16) and outperforms existing state-of-the-art methods across various LLMs and datasets. Notably, CBQ only takes 4.3 hours to quantize a weightonly quantization of a 4-bit LLAMA1-65B model, achieving a commendable trade off between performance and efficiency.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Q-VLM: Post-training Quantization for Large Vision-Language ModelsChangyuan Wang, Ziwei Wang, Xiuwei Xu, Yansong Tang 等NeurIPS 2024 · 被引用 57 次
- Model-Preserving Adaptive RoundingAlbert Tseng, Zhaofeng Sun, Chris De SaICML 2026 · 被引用 17 次
- SliderQuant: Accurate Post-Training Quantization for LLMsShigeng Wang, Chao Li, Yangyuxuan Kang, Jiawei Fan 等ICLR 2026 · 被引用 6 次
- Learning Grouped Lattice Vector Quantizers for Low-Bit LLM CompressionXi Zhang, Xiaolin Wu, Jiamang Wang, Weisi LinNeurIPS 2025 · 被引用 4 次
- Task-Specific Zero-Shot Quantization-Aware Training for Object DetectionChanghao Li, Xinrui Chen, Ji Wang, Kang Zhao 等ICCV 2025 · 被引用 2 次
它引用的顶会 Paper21
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language ModelsGuangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu 等ICML 2023 · 被引用 1,493 次
- Aligning AI With Shared Human ValuesDan Hendrycks, Collin Burns, Steven Basart, Andrew Critch 等ICLR 2021 · 被引用 878 次
相关 Paper
- ACBQ: Adaptive Cross-Block Quantization of Large Language ModelsHailing Wang, Jianglin Lu, Yitian Zhang, Huimin Zeng 等ACL 2026
- SERQ: Saliency-Aware Low-Rank Error Reconstruction for LLM QuantizationYeonsik Park, Hyeonseong Kim, Seungkyu ChoiICLR 2026 · 被引用 1 次
- Layer-Wise High-Impact Parameter Ratio Optimization in Post-Training Quantization for Large Language ModelsCuong Pham, Dung Anh Hoang, Cuong C. Nguyen, Trung Le 等ACL 2026
- PTQ1.61: Push the Real Limit of Extremely Low-Bit Post-Training Quantization Methods for Large Language ModelsJiaqi Zhao, Miao Zhang, Ming Wang, Yuzhang Shang 等ACL 2025 · 被引用 7 次
- QLLM: Accurate and Efficient Low-Bitwidth Quantization for Large Language ModelsJing Liu, Ruihao Gong, Xiuying Wei, Zhiwei Dong 等ICLR 2024 · 被引用 75 次
