Confidence Under the Hood: An Investigation into the Confidence-Probability Alignment in Large Language Models
Abhishek Kumar, Robert Morabito, Sanzhar Umbet, Jad Kabbara, Ali Emami
摘要
As the use of Large Language Models (LLMs) becomes more widespread, understanding their self-evaluation of confidence in generated responses becomes increasingly important as it is integral to the reliability of the output of these models. We introduce the concept of Confidence-Probability Alignment, that connects an LLM's internal confidence, quantified by token probabilities, to the confidence conveyed in the model's response when explicitly asked about its certainty. Using various datasets and prompting techniques that encourage model introspection, we probe the alignment between models' internal and expressed confidence. These techniques encompass using structured evaluation scales to rate confidence, including answer options when prompting, and eliciting the model's confidence level for outputs it does not recognize as its own. Notably, among the models analyzed, OpenAI's GPT-4 showed the strongest confidence-probability alignment, with an average Spearman's ρ of 0.42, across a wide range of tasks. Our work contributes to the ongoing efforts to facilitate risk assessment in the application of LLMs and to further our understanding of model trustworthiness. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Generate, but Verify: Reducing Hallucination in Vision-Language Models with Retrospective ResamplingTsung-Han Wu, Heekyung Lee, Jiaxin Ge, Joseph E. Gonzalez 等NeurIPS 2025 · 被引用 33 次
- DTS: Enhancing Large Reasoning Models via Decoding Tree SketchingZicheng Xu, Xiuyi Lou, Guanchu Wang, Yu-Neng Chuang 等ICML 2026 · 被引用 5 次
- On the Robustness of Verbal Confidence of LLMs in Adversarial AttacksStephen Obadinma, Xiaodan ZhuNeurIPS 2025 · 被引用 3 次
- DDO: Dual-Decision Optimization for LLM-Based Medical Consultation via Multi-Agent CollaborationZhihao Jia, Mingyi Jia, Junwen Duan, Jian-xin WangEMNLP 2025 · 被引用 2 次
- Completing Missing Annotation: Multi-Agent Debate for Accurate and Scalable Relevant Assessment for IR BenchmarksMinjeong Ban, Jeonghwan Choi, Hyangsuk Min, Nicole Hee-Yeon Kim 等ICLR 2026 · 被引用 1 次
它引用的顶会 Paper11
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
- Improving Factuality and Reasoning in Language Models through Multiagent DebateYilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum 等ICML 2024 · 被引用 1,562 次
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le 等ICLR 2023 · 被引用 681 次
相关 Paper
- Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMsMiao Xiong, Zhiyuan Hu, Xinyang Lu, Yifei Li 等ICLR 2024 · 被引用 867 次
- Learning to Route LLMs with Confidence TokensYu-Neng Chuang, Prathusha Kameswara Sarma, Parikshit Gopalan, John Boccio 等ICML 2025
- From Sampling to Cognition: Modeling Internal Cognitive Confidence in Language Models for Robust Uncertainty CalibrationHao Li, Tao He, Jiafeng Liang, Zheng Chu 等AAAI 2026
- Multicalibration for Confidence Scoring in LLMsGianluca Detommaso, Martin Bertran Lopez, Riccardo Fogliato, Aaron RothICML 2024 · 被引用 39 次
- MetaFaith: Faithful Natural Language Uncertainty Expression in LLMsGabrielle Kaili-May Liu, Gal Yona, Avi Caciularu, Idan Szpektor 等EMNLP 2025
