Quantized Can Still Be Calibrated: A Unified Framework to Calibration in Quantized Large Language Models
Mingyu Zhong, Guanchu Wang, Yu-Neng Chuang, Na Zou
摘要
Although weight quantization helps large language models (LLMs) in resource-constrained environments, its influence on the uncertainty calibration remains unexplored. To bridge this gap, we present a comprehensive investigation of uncertainty calibration for quantized LLMs in this work. Specifically, we propose an analytic method to estimate the upper bound of calibration error (UBCE) for LLMs. Our method separately discusses the calibration error of the model's correct and incorrect predictions, indicating a theoretical improvement of calibration error caused by weight quantization. Our study demonstrates that quantized models consistently exhibit worse calibration performance than full-precision models, supported by consistent analysis across multiple LLMs and datasets. To address the calibration issues of quantized models, we propose a novel post-calibration method to recover the calibration performance of quantized models through soft-prompt tuning. Specifically, we inject soft tokens into quantized models after the embedding layers and optimize these tokens to recover the calibration error caused by weight quantization. Experimental results on multiple datasets demonstrate its effectiveness in improving the uncertainty calibration of quantized LLMs, facilitating more reliable weight quantization in resource-constrained environments.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper2
相关 Paper
- EasyQuant: An Efficient Data-free Quantization Algorithm for LLMsHanlin Tang, Yifu Sun, Decheng Wu, Kai Liu 等EMNLP 2023 · 被引用 4 次
- Assessing Safety Risks and Quantization-aware Safety Patching for Quantized Large Language ModelsKejia Chen, Jiawen Zhang, Jiacong Hu, Yu Wang 等ICML 2025
- Low-Bit Quantization Favors Undertrained LLMsXu Ouyang, Tao Ge, Thomas Hartvigsen, Zhisong Zhang 等ACL 2025 · 被引用 3 次
- FAIR-Calib: Frontier-Aware Instability-Reweighted Calibration for Post-Training Quantization of Diffusion Large Language ModelsHaoyu Huang, Linlin Yang, Sheng Xu, Boyu Liu 等ICML 2026
- SliderQuant: Accurate Post-Training Quantization for LLMsShigeng Wang, Chao Li, Yangyuxuan Kang, Jiawei Fan 等ICLR 2026 · 被引用 6 次
