Rewarding Doubt: A Reinforcement Learning Approach to Calibrated Confidence Expression of Large Language Models
David Bani-Harouni, Chantal Pellegrini, Paul Stangel, Ege Özsoy, Kamilia Zaripova, Nassir Navab, Matthias Keicher
Abstract
A safe and trustworthy use of Large Language Models (LLMs) requires an accurate expression of confidence in their answers. We propose a novel Reinforcement Learning approach that allows to directly fine-tune LLMs to express calibrated confidence estimates alongside their answers to factual questions. Our method optimizes a reward based on the logarithmic scoring rule, explicitly penalizing both over- and under-confidence. This encourages the model to align its confidence estimates with the actual predictive accuracy. The optimal policy under our reward design would result in perfectly calibrated confidence expressions. Unlike prior approaches that decouple confidence estimation from response generation, our method integrates confidence calibration seamlessly into the generative process of the LLM. Empirically, we demonstrate that models trained with our approach exhibit substantially improved calibration and generalize to unseen tasks without further fine-tuning, suggesting the emergence of general confidence awareness. Our code is available at https://github.com/pasta99/RewardingDoubt.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1b18d960-0eb6-45da-9702-7b391c97a712Cited by top-tier papers7
- BARREL: Boundary-Aware Reasoning for Factual and Reliable LRMsJunxiao Yang, Jinzhe Tu, Haoran Liu, Xiaoce Wang et al.ICLR 2026 · 9 citations
- VL-Calibration: Decoupled Confidence Calibration for Large Vision-Language Models ReasoningWenyi Xiao, Xinchi Xu, Leilei GanACL 2026 · 2 citations
- Deliberative Searcher: Improving LLM Reliability via Reinforcement Learning with ConstraintsZhenyun Yin, Shujie Wang, Xuhong Wang, Xingjun Ma et al.ACL 2026 · 1 citation
- Auto-Fill: Learning to Predict Missing Values Accurately with Specialist Language ModelsYurong Liu, Yeye He, Haoyu Dong, Junjie Xing et al.VLDB 2026
- Miner: Mining Intrinsic Mastery for Data-Efficient RL in Large Reasoning ModelsShuyang Jiang, Yuhao Wang, Ya Zhang, Yanfeng Wang et al.ACL 2026
Builds on13
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMsMiao Xiong, Zhiyuan Hu, Xinyang Lu, Yifei Li et al.ICLR 2024 · 867 citations
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le et al.ICLR 2023 · 681 citations
Related papers
- Linguistic Calibration of Long-Form GenerationsNeil Band, Xuechen Li, Tengyu Ma, Tatsunori HashimotoICML 2024 · 56 citations
- The Confidence Dichotomy: Analyzing and Mitigating Miscalibration in Tool-Use AgentsWeihao Xuan, Qingcheng Zeng, Heli Qi, Yunze Xiao et al.ACL 2026 · 4 citations
- Beyond Binary Rewards: Training LMs to Reason About Their UncertaintyMehul Damani, Isha Puri, Stewart Slocum, Idan Shenfeld et al.ICLR 2026 · 116 citations
- ADVICE: Answer-Dependent Verbalized Confidence EstimationKi Jung Seo, Sehun Lim, Taeuk KimACL 2026 · 4 citations
- ConfTuner: Training Large Language Models to Express Their Confidence VerballyYibo Li, Miao Xiong, Jiaying Wu, Bryan HooiNeurIPS 2025 · 43 citations
