Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty
Mehul Damani, Isha Puri, Stewart Slocum, Idan Shenfeld, Leshem Choshen, Yoon Kim, Jacob Andreas
Abstract
When language models (LMs) are trained via reinforcement learning (RL) to generate natural language reasoning chains, their performance improves on a variety of difficult question answering tasks. Today, almost all successful applications of RL for reasoning use binary reward functions that evaluate the correctness of LM outputs. Because such reward functions do not penalize guessing or low-confidence outputs, they often have the unintended side-effect of degrading calibration and increasing the rate at which LMs generate incorrect responses (i.e. "hallucinate ′′ ) in other problem domains. This paper describes RLCR (Reinforcement Learning with Calibration Rewards), an approach to training reasoning models that jointly improves accuracy and calibrated confidence estimation. During RLCR, LMs generate both predictions and numerical confidence estimates after reasoning. They are trained to optimize a reward function that augments a binary correctness score with a Brier score-a scoring rule for confidence estimates that incentivizes calibrated prediction. We first prove that this reward function (or any analogous reward function that uses a bounded, proper scoring rule) yields models whose predictions are both accurate and well-calibrated. We next show that across diverse datasets, RLCR substantially improves calibration while maintaining strong accuracy on both in-domain and out-of-domain evaluations-outperforming both ordinary RL training and classifiers trained to assign post-hoc confidence scores. While ordinary RL hurts calibration, RLCR improves it. Finally, we demonstrate that verbalized confidence can be leveraged at test time to improve accuracy and calibration via confidence-weighted scaling methods. Our results show that explicitly optimizing for calibration can produce more generally reliable reasoning models. Code, models, demonstrations and further information is available at https://rl-calibration.github.io/
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f9ef8020-588a-4dfc-b305-98604527ffdfCited by top-tier papers23
- Time Is a Feature: Exploiting Temporal Dynamics in Diffusion Language ModelsWen Wang, Bozhen Fang, Chenchen Jing, Yongliang Shen et al.ICLR 2026 · 33 citations
- BrokenMath: A Benchmark for Sycophancy in Theorem Proving with LLMsIvo Petrov, Jasper Dekoninck, Martin VechevICML 2026 · 25 citations
- BARREL: Boundary-Aware Reasoning for Factual and Reliable LRMsJunxiao Yang, Jinzhe Tu, Haoran Liu, Xiaoce Wang et al.ICLR 2026 · 9 citations
- Zero-Overhead Introspection for Adaptive Test-Time ComputeRohin Manvi, Joey Hong, Tim Seyde, Maxime Labonne et al.ICLR 2026 · 7 citations
- Generalized Correctness Models: Learning Calibrated and Cross-Model Correctness Predictors from Historical PatternsHanqi Xiao, Vaidehi Patil, Hyunji Lee, Elias Stengel-Eskin et al.ICML 2026 · 5 citations
Builds on14
- Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMsMiao Xiong, Zhiyuan Hu, Xinyang Lu, Yifei Li et al.ICLR 2024 · 867 citations
- Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base ModelJingcheng Hu, Yinmin Zhang, Qi Han, Daxin Jiang et al.NeurIPS 2025 · 533 citations
- Scalable Best-of-N Selection for Large Language Models via Self-CertaintyZhewei Kang, Xuandong Zhao, Dawn SongNeurIPS 2025 · 211 citations
- RL's Razor: Why Online Reinforcement Learning Forgets LessIdan Shenfeld, Jyothish Pari, Pulkit AgrawalICLR 2026 · 176 citations
- On the Weaknesses of Reinforcement Learning for Neural Machine TranslationLeshem Choshen, Lior Fox, Zohar Aizenbud, Omri AbendICLR 2020 · 124 citations
Related papers
- Rewarding Doubt: A Reinforcement Learning Approach to Calibrated Confidence Expression of Large Language ModelsDavid Bani-Harouni, Chantal Pellegrini, Paul Stangel, Ege Özsoy et al.ICLR 2026 · 49 citations
- VL-Calibration: Decoupled Confidence Calibration for Large Vision-Language Models ReasoningWenyi Xiao, Xinchi Xu, Leilei GanACL 2026 · 2 citations
- Linguistic Calibration of Long-Form GenerationsNeil Band, Xuechen Li, Tengyu Ma, Tatsunori HashimotoICML 2024 · 56 citations
- Drift: Enhancing LLM Faithfulness in Rationale Generation via Dual-Reward Probabilistic InferenceJiazheng Li, Hanqi Yan, Yulan HeACL 2025
- ConfTuner: Training Large Language Models to Express Their Confidence VerballyYibo Li, Miao Xiong, Jiaying Wu, Bryan HooiNeurIPS 2025 · 43 citations
