Restoring Calibration for Aligned Large Language Models: A Calibration-Aware Fine-Tuning Approach
Jiancong Xiao, Bojian Hou, Zhanliang Wang, Ruochen Jin, Qi Long, Weijie J. Su, Li Shen
Abstract
One of the key technologies for the success of Large Language Models (LLMs) is preference alignment. However, a notable side effect of preference alignment is poor calibration: while the pre-trained models are typically well-calibrated, LLMs tend to become poorly calibrated after alignment with human preferences. In this paper, we investigate why preference alignment affects calibration and how to address this issue. For the first question, we observe that the preference collapse issue in alignment undesirably generalizes to the calibration scenario, causing LLMs to exhibit overconfidence and poor calibration. To address this, we demonstrate the importance of fine-tuning with domain-specific knowledge to alleviate the overconfidence issue. To further analyze whether this affects the model's performance, we categorize models into two regimes: calibratable and non-calibratable, defined by bounds of Expected Calibration Error (ECE). In the calibratable regime, we propose a calibration-aware fine-tuning approach to achieve proper calibration without compromising LLMs' performance. However, as models are further finetuned for better performance, they enter the noncalibratable regime. For this case, we develop an EM-algorithm-based ECE regularization for the fine-tuning loss to maintain low calibration error. Extensive experiments validate the effectiveness of the proposed methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e221d976-d528-4150-95bd-4e1503252691Cited by top-tier papers5
- Trained on Tokens, Calibrated on Concepts: The Emergence of Semantic Calibration in LLMsPreetum Nakkiran, Arwen Bradley, Adam Golinski, Eugène Ndiaye et al.ICLR 2026 · 17 citations
- BaseCal: Unsupervised Confidence Calibration via Base Model SignalsHexiang Tan, Wanli Yang, Junwei Zhang, Xin Chen et al.ACL 2026 · 3 citations
- Calibration-Aware Policy Optimization for Reasoning LLMsZiqi Wang, Xingzhou Lou, Meiqi Wu, Zhengqi Wen et al.ACL 2026 · 2 citations
- IA2: Alignment with ICL Activations improves Supervised Fine-TuningAayush Mishra, Daniel Khashabi, Anqi LiuICLR 2026 · 1 citation
- The Geometry of Narrow Fine-Tuning Degradation: Trajectory Lock-in and Spectral BifurcationJia Liu, Jiaxin Luo, Xinhao Qiu, Yixue Hao et al.ICML 2026
Builds on26
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Calibrate Before Use: Improving Few-shot Performance of Language ModelsZihao Zhao, Eric Wallace, Shi Feng, Dan Klein et al.ICML 2021 · 1,843 citations
- Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMsMiao Xiong, Zhiyuan Hu, Xinyang Lu, Yifei Li et al.ICLR 2024 · 867 citations
- Calibrating Deep Neural Networks using Focal LossJishnu Mukhoti, Viveka Kulharia, Amartya Sanyal, Stuart Golodetz et al.NeurIPS 2020 · 674 citations
- Statistical Rejection Sampling Improves Preference OptimizationTianqi Liu, Yao Zhao, Rishabh Joshi, Misha Khalman et al.ICLR 2024 · 346 citations
Related papers
- Towards Objective Fine-tuning: How LLMs' Prior Knowledge Causes Potential Poor Calibration?Ziming Wang, Zeyu Shi, Haoyi Zhou, Shiqi Gao et al.ACL 2025 · 6 citations
- Leveraging robust optimization for llm alignment under distribution shiftsMingye Zhu, Yi Liu, Zheren Fu, Yongdong Zhang et al.NeurIPS 2025 · 2 citations
- Provably Mitigating Overoptimization in RLHF: Your SFT Loss is Implicitly an Adversarial RegularizerZhihan Liu, Miao Lu, Shenao Zhang, Boyi Liu et al.NeurIPS 2024 · 119 citations
- Preserving Pre-trained Features Helps Calibrate Fine-tuned Language ModelsGuande He, Jianfei Chen, Jun ZhuICLR 2023 · 1 citation
- Preference-Oriented Supervised Fine-Tuning: Favoring Target Model over Aligned Large Language ModelsYuchen Fan, Yuzhong Hong, Qiushi Wang, Junwei Bao et al.AAAI 2025 · 7 citations
