Lune

ACL2025Top-tier venue

Quantized Can Still Be Calibrated: A Unified Framework to Calibration in Quantized Large Language Models

Mingyu Zhong, Guanchu Wang, Yu-Neng Chuang, Na Zou

2025Year

Abstract

Although weight quantization helps large language models (LLMs) in resource-constrained environments, its influence on the uncertainty calibration remains unexplored. To bridge this gap, we present a comprehensive investigation of uncertainty calibration for quantized LLMs in this work. Specifically, we propose an analytic method to estimate the upper bound of calibration error (UBCE) for LLMs. Our method separately discusses the calibration error of the model's correct and incorrect predictions, indicating a theoretical improvement of calibration error caused by weight quantization. Our study demonstrates that quantized models consistently exhibit worse calibration performance than full-precision models, supported by consistent analysis across multiple LLMs and datasets. To address the calibration issues of quantized models, we propose a novel post-calibration method to recover the calibration performance of quantized models through soft-prompt tuning. Specifically, we inject soft tokens into quantized models after the embedding layers and optimize these tokens to recover the calibration error caused by weight quantization. Experimental results on multiple datasets demonstrate its effectiveness in improving the uncertainty calibration of quantized LLMs, facilitating more reliable weight quantization in resource-constrained environments.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

Builds on2

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines