Enhancing Uncertainty Estimation in LLMs with Expectation of Aggregated Internal Belief
Zeguan Xiao, Diyang Dou, Boya Xiong, Yun Chen, Guanhua Chen
Abstract
Large Language Models (LLMs) have achieved remarkable success across a wide range of natural language tasks, but often exhibit overconfidence and generate plausible yet incorrect answers. This overconfidence, especially in models undergone Reinforcement Learning from Human Feedback (RLHF), poses significant challenges for reliable uncertainty estimation and safe deployment. In this paper, we propose EAGLE (Expectation of AGgregated internaL bEief), a novel self-evaluation-based calibration method that leverages the internal hidden states of LLMs to derive more accurate confidence scores. Instead of relying on the model's final output, our approach extracts internal beliefs from multiple intermediate layers during self-evaluation. By aggregating these layer-wise beliefs and calculating the expectation over the resulting confidence score distribution, EAGLE produces a refined confidence score that more faithfully reflects the model's internal certainty. Extensive experiments on diverse datasets and LLMs demonstrate that EAGLE significantly improves calibration performance over existing baselines. We also provide an in-depth analysis of EAGLE, including a layer-wise examination of uncertainty patterns, a study of the impact of self-evaluation prompts, and an analysis of the effect of self-evaluation score range. Our code is public at https://github.com/sustech-nlp/EAGLE .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Dynamics Within Latent Chain-of-Thought: An Empirical Study of Causal StructureZirui Li, Xuefeng Bai, Kehai Chen, Yizhi Li et al.ICML 2026
- Code-MUE: Measuring Code LLMs’ Uncertainty through Execution-Based Semantic Interaction GraphsXiaoning Ren, Yinxing Xue, Lei Ma, Yuheng HuangISSTA 2026
Builds on9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMsMiao Xiong, Zhiyuan Hu, Xinyang Lu, Yifei Li et al.ICLR 2024 · 867 citations
- Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary SpaceMor Geva, Avi Caciularu, Kevin Ro Wang, Yoav GoldbergEMNLP 2022 · 92 citations
- Discovering Latent Knowledge in Language Models Without SupervisionCollin Burns, Haotian Ye, Dan Klein, Jacob SteinhardtICLR 2023 · 45 citations
Related papers
- Calibrating the Confidence of Large Language Models by Eliciting FidelityMozhi Zhang, Mianqiu Huang, Rundong Shi, Linsen Guo et al.EMNLP 2024 · 5 citations
- Rewarding Doubt: A Reinforcement Learning Approach to Calibrated Confidence Expression of Large Language ModelsDavid Bani-Harouni, Chantal Pellegrini, Paul Stangel, Ege Özsoy et al.ICLR 2026 · 49 citations
- Multicalibration for Confidence Scoring in LLMsGianluca Detommaso, Martin Bertran Lopez, Riccardo Fogliato, Aaron RothICML 2024 · 39 citations
- From Sampling to Cognition: Modeling Internal Cognitive Confidence in Language Models for Robust Uncertainty CalibrationHao Li, Tao He, Jiafeng Liang, Zheng Chu et al.AAAI 2026
- SaySelf: Teaching LLMs to Express Confidence with Self-Reflective RationalesTianyang Xu, Shujin Wu, Shizhe Diao, Xiaoze Liu et al.EMNLP 2024 · 10 citations
