Uncertainty in Language Models: Assessment through Rank-Calibration
Xinmeng Huang, Shuo Li, Mengxin Yu, Matteo Sesia, Hamed Hassani, Insup Lee, Osbert Bastani, Edgar Dobriban
Abstract
Language Models (LMs) have shown promising performance in natural language generation. However, as LMs often generate incorrect or hallucinated responses, it is crucial to correctly quantify their uncertainty in responding to given inputs. In addition to verbalized confidence elicited via prompting, many uncertainty measures (e.g., semantic entropy and affinitygraph-based measures) have been proposed. However, these measures can differ greatly, and it is unclear how to compare them, partly because they take values over different ranges (e.g., [0, ∞) or [0, 1]). In this work, we address this issue by developing a novel and practical framework, termed Rank-Calibration, to assess uncertainty and confidence measures for LMs. Our key tenet is that higher uncertainty (or lower confidence) should imply lower generation quality, on average. Rank-calibration quantifies deviations from this ideal relationship in a principled manner, without requiring ad hoc binary thresholding of the correctness score (e.g., ROUGE or METEOR). The broad applicability and the granular interpretability of our methods are demonstrated empirically. The code to replicate our experiments is here.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4b654cec-7554-4e6d-80b0-63b110993578Cited by top-tier papers8
- Conformal Alignment: Knowing When to Trust Foundation Models with GuaranteesYu Gui, Ying Jin, Zhimei RenNeurIPS 2024 · 63 citations
- One-Shot Safety Alignment for Large Language Models via Optimal DualizationXinmeng Huang, Shuo Li, Edgar Dobriban, Osbert Bastani et al.NeurIPS 2024 · 30 citations
- MUR: Momentum Uncertainty guided Reasoning for Large Language ModelsHang Yan, Fangzhi Xu, Rongman Xu, Yifei Li et al.ACL 2026 · 12 citations
- Can LLMs Refuse Questions They Do Not Know? Measuring Knowledge-Aware Refusal in Factual TasksWenbo Pan, Jie Xu, Qiguang Chen, Junhao Dong et al.ICLR 2026 · 8 citations
- Temporal Difference Calibration in Sequential Tasks: Application to Vision-Language-Action ModelsShelly Francis-Meretzki, Mirco Mutti, Yaniv Romano, Aviv TamarICML 2026 · 2 citations
Builds on5
- Calibrate Before Use: Improving Few-shot Performance of Language ModelsZihao Zhao, Eric Wallace, Shi Feng, Dan Klein et al.ICML 2021 · 1,843 citations
- Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMsMiao Xiong, Zhiyuan Hu, Xinyang Lu, Yifei Li et al.ICLR 2024 · 867 citations
- SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language ModelsPotsawee Manakul, Adian Liusie, Mark J. F. GalesEMNLP 2023 · 331 citations
- Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language GenerationLorenz Kuhn, Yarin Gal, Sebastian FarquharICLR 2023 · 49 citations
- LUQ: Long-text Uncertainty Quantification for LLMsCaiqi Zhang, Fangyu Liu, Marco Basaldella, Nigel CollierEMNLP 2024 · 11 citations
Related papers
- Addressing Pitfalls in the Evaluation of Uncertainty Estimation Methods for Natural Language GenerationMykyta Ielanskyi, Kajetan Schweighofer, Lukas Aichberger, Sepp HochreiterICLR 2026 · 10 citations
- Calibrating the Confidence of Large Language Models by Eliciting FidelityMozhi Zhang, Mianqiu Huang, Rundong Shi, Linsen Guo et al.EMNLP 2024 · 5 citations
- Inv-Entropy: A Fully Probabilistic Framework for Uncertainty Quantification in Language ModelsHaoyi Song, Ruihan Ji, Naichen Shi, Fan Lai et al.NeurIPS 2025 · 6 citations
- Kernel Language Entropy: Fine-grained Uncertainty Quantification for LLMs from Semantic SimilaritiesAlexander Nikitin, Jannik Kossen, Yarin Gal, Pekka MarttinenNeurIPS 2024 · 197 citations
- Mapping from Meaning: Addressing the Miscalibration of Prompt-Sensitive Language ModelsKyle Cox, Jiawei Xu, Yikun Han, Rong Xu et al.AAAI 2025 · 6 citations
