Taming Overconfidence in LLMs: Reward Calibration in RLHF
Jixuan Leng, Chengsong Huang, Banghua Zhu, Jiaxin Huang
Abstract
Language model calibration refers to the alignment between the confidence of the model and the actual performance of its responses. While previous studies point out the overconfidence phenomenon in Large Language Models (LLMs) and show that LLMs trained with Reinforcement Learning from Human Feedback (RLHF) are overconfident with a more sharpened output probability, in this study, we reveal that RLHF tends to lead models to express verbalized overconfidence in their own responses. We investigate the underlying cause of this overconfidence and demonstrate that reward models used for Proximal Policy Optimization (PPO) exhibit inherent biases towards high-confidence scores regardless of the actual quality of responses. Building upon this insight, we propose two PPO variants: PPO-M: PPO with Calibrated Reward Modeling and PPO-C: PPO with Calibrated Reward Calculation. PPO-M integrates explicit confidence scores in reward model training, which calibrates reward models to better capture the alignment between response quality and verbalized confidence. PPO-C adjusts the reward score during PPO based on the difference between the current reward and the exponential average of past rewards. Both PPO-M and PPO-C can be seamlessly integrated into the current PPO pipeline and do not require additional golden labels. We evaluate our methods on both Llama3-8B and Mistral-7B across six diverse datasets including multiple-choice and open-ended generation. Experimental results demonstrate that both of our methods can reduce calibration error and maintain performance comparable to standard PPO. We further show that they could preserve model capabilities in open-ended conversational settings. Our code is publicly released. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers29
- Beyond Binary Rewards: Training LMs to Reason About Their UncertaintyMehul Damani, Isha Puri, Stewart Slocum, Idan Shenfeld et al.ICLR 2026 · 116 citations
- Rewarding Doubt: A Reinforcement Learning Approach to Calibrated Confidence Expression of Large Language ModelsDavid Bani-Harouni, Chantal Pellegrini, Paul Stangel, Ege Özsoy et al.ICLR 2026 · 49 citations
- Trained on Tokens, Calibrated on Concepts: The Emergence of Semantic Calibration in LLMsPreetum Nakkiran, Arwen Bradley, Adam Golinski, Eugène Ndiaye et al.ICLR 2026 · 17 citations
- Calibrating Verbalized Confidence with Self-Generated DistractorsVictor Wang, Elias Stengel-EskinICLR 2026 · 15 citations
- Shrinking the Variance: Shrinkage Baselines for Reinforcement Learning with Verifiable RewardsGuanning Zeng, Zhaoyi Zhou, Daman Arora, Andrea ZanetteICML 2026 · 13 citations
Builds on20
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
Related papers
- Decoupling Reasoning and Confidence: Resurrecting Calibration in Reinforcement Learning from Verifiable RewardsZhengzhao Ma, Xueru Wen, Boxi Cao, Yaojie Lu et al.ICML 2026 · 5 citations
- Don't Forget Your Reward Values: Language Model Alignment via Value-based CalibrationXin Mao, Feng-Lin Li, Huimin Xu, Wei Zhang et al.EMNLP 2024 · 1 citation
- OPPO: Accelerating PPO-based RLHF via Pipeline OverlapKaizhuo Yan, Yingjie Yu, Yifan Yu, Haizhong Zheng et al.ICLR 2026 · 4 citations
- ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language ModelsZiniu Li, Tian Xu, Yushun Zhang, Zhihang Lin et al.ICML 2024 · 165 citations
- TUR-DPO: Topology- and Uncertainty-Aware Direct Preference OptimizationAbdulhady abas, Fatemeh Daneshfar, Seyedali Mirjalili, Mourad OussalahICML 2026
