Addressing Pitfalls in the Evaluation of Uncertainty Estimation Methods for Natural Language Generation
Mykyta Ielanskyi, Kajetan Schweighofer, Lukas Aichberger, Sepp Hochreiter
Abstract
Hallucinations are a common issue that undermine the reliability of large language models (LLMs). Recent studies have identified a specific subset of hallucinations, known as confabulations, which arise due to predictive uncertainty of LLMs. To detect confabulations, various methods for estimating predictive uncertainty in natural language generation (NLG) have been developed. These methods are typically evaluated by correlating uncertainty estimates with the correctness of generated text, with question-answering (QA) datasets serving as the standard benchmark. However, commonly used approximate correctness functions have substantial disagreement between each other and, consequently, in the ranking of the uncertainty estimation methods. This allows one to inflate the apparent performance of uncertainty estimation methods. We propose using several alternative risk indicators for risk correlation experiments that improve robustness of empirical assessment of uncertainty estimation algorithms for NLG. For QA tasks, we show that marginalizing over multiple LLM-as-a-judge variants leads to reducing the evaluation biases. Furthermore, we explore structured tasks as well as out of distribution and perturbation detection tasks which provide robust and controllable risk indicators. Finally, we propose to use an Elo rating of uncertainty estimation methods to give an objective summarization over extensive evaluation settings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 33f8eccd-f4a7-45cc-8154-970c61b99d33Cited by top-tier papers2
- Complementing Self-Consistency with Cross-Model Disagreement for Uncertainty QuantificationKimia Hamidieh, Veronika Thost, Walter Gerych, Mikhail Yurochkin et al.ICLR 2026 · 12 citations
- UCPO: Uncertainty-Aware Policy OptimizationXianzhou Zeng, Jing Huang, Chunmei Xie, Gongrui Nan et al.ICML 2026
Builds on22
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos et al.ICML 2024 · 1,212 citations
- Bayesian Deep Learning and a Probabilistic Perspective of GeneralizationAndrew Gordon Wilson, Pavel IzmailovNeurIPS 2020 · 845 citations
- Laplace Redux - Effortless Bayesian Deep LearningErik A. Daxberger, Agustinus Kristiadi, Alexander Immer, Runa Eschenhagen et al.NeurIPS 2021 · 508 citations
- Uncertainty Estimation in Autoregressive Structured PredictionAndrey Malinin, Mark J. F. GalesICLR 2021 · 439 citations
- SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language ModelsPotsawee Manakul, Adian Liusie, Mark J. F. GalesEMNLP 2023 · 331 citations
Related papers
- Improving Uncertainty Estimation through Semantically Diverse Language GenerationLukas Aichberger, Kajetan Schweighofer, Mykyta Ielanskyi, Sepp HochreiterICLR 2025
- Probabilities Are All You Need: A Probability-Only Approach to Uncertainty Estimation in Large Language ModelsManh Nguyen, Sunil Gupta, Hung LeAAAI 2026 · 4 citations
- Estimating LLM Consistency: A User Baseline vs Surrogate MetricsXiaoyuan Wu, Weiran Lin, Omer Akgul, Lujo BauerEMNLP 2025
- Can We Trust the Performance Evaluation of Uncertainty Estimation Methods in Text Summarization?Jianfeng He, Runing Yang, Linlin Yu, Changbin Li et al.EMNLP 2024 · 1 citation
- To Believe or Not to Believe Your LLM: Iterative Prompting for Estimating Epistemic UncertaintyYasin Abbasi-Yadkori, Ilja Kuzborskij, András György, Csaba SzepesváriNeurIPS 2024
