Lune

ACL2025顶会

Quantized Can Still Be Calibrated: A Unified Framework to Calibration in Quantized Large Language Models

Mingyu Zhong, Guanchu Wang, Yu-Neng Chuang, Na Zou

2025年份

摘要

Although weight quantization helps large language models (LLMs) in resource-constrained environments, its influence on the uncertainty calibration remains unexplored. To bridge this gap, we present a comprehensive investigation of uncertainty calibration for quantized LLMs in this work. Specifically, we propose an analytic method to estimate the upper bound of calibration error (UBCE) for LLMs. Our method separately discusses the calibration error of the model's correct and incorrect predictions, indicating a theoretical improvement of calibration error caused by weight quantization. Our study demonstrates that quantized models consistently exhibit worse calibration performance than full-precision models, supported by consistent analysis across multiple LLMs and datasets. To address the calibration issues of quantized models, we propose a novel post-calibration method to recover the calibration performance of quantized models through soft-prompt tuning. Specifically, we inject soft tokens into quantized models after the embedding layers and optimize these tokens to recover the calibration error caused by weight quantization. Experimental results on multiple datasets demonstrate its effectiveness in improving the uncertainty calibration of quantized LLMs, facilitating more reliable weight quantization in resource-constrained environments.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper2

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖