Uncertainty in Language Models: Assessment through Rank-Calibration
Xinmeng Huang, Shuo Li, Mengxin Yu, Matteo Sesia, Hamed Hassani, Insup Lee, Osbert Bastani, Edgar Dobriban
摘要
Language Models (LMs) have shown promising performance in natural language generation. However, as LMs often generate incorrect or hallucinated responses, it is crucial to correctly quantify their uncertainty in responding to given inputs. In addition to verbalized confidence elicited via prompting, many uncertainty measures (e.g., semantic entropy and affinitygraph-based measures) have been proposed. However, these measures can differ greatly, and it is unclear how to compare them, partly because they take values over different ranges (e.g., [0, ∞) or [0, 1]). In this work, we address this issue by developing a novel and practical framework, termed Rank-Calibration, to assess uncertainty and confidence measures for LMs. Our key tenet is that higher uncertainty (or lower confidence) should imply lower generation quality, on average. Rank-calibration quantifies deviations from this ideal relationship in a principled manner, without requiring ad hoc binary thresholding of the correctness score (e.g., ROUGE or METEOR). The broad applicability and the granular interpretability of our methods are demonstrated empirically. The code to replicate our experiments is here.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Conformal Alignment: Knowing When to Trust Foundation Models with GuaranteesYu Gui, Ying Jin, Zhimei RenNeurIPS 2024 · 被引用 63 次
- One-Shot Safety Alignment for Large Language Models via Optimal DualizationXinmeng Huang, Shuo Li, Edgar Dobriban, Osbert Bastani 等NeurIPS 2024 · 被引用 30 次
- MUR: Momentum Uncertainty guided Reasoning for Large Language ModelsHang Yan, Fangzhi Xu, Rongman Xu, Yifei Li 等ACL 2026 · 被引用 12 次
- Can LLMs Refuse Questions They Do Not Know? Measuring Knowledge-Aware Refusal in Factual TasksWenbo Pan, Jie Xu, Qiguang Chen, Junhao Dong 等ICLR 2026 · 被引用 8 次
- Temporal Difference Calibration in Sequential Tasks: Application to Vision-Language-Action ModelsShelly Francis-Meretzki, Mirco Mutti, Yaniv Romano, Aviv TamarICML 2026 · 被引用 2 次
它引用的顶会 Paper5
- Calibrate Before Use: Improving Few-shot Performance of Language ModelsZihao Zhao, Eric Wallace, Shi Feng, Dan Klein 等ICML 2021 · 被引用 1,843 次
- Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMsMiao Xiong, Zhiyuan Hu, Xinyang Lu, Yifei Li 等ICLR 2024 · 被引用 867 次
- SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language ModelsPotsawee Manakul, Adian Liusie, Mark J. F. GalesEMNLP 2023 · 被引用 331 次
- Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language GenerationLorenz Kuhn, Yarin Gal, Sebastian FarquharICLR 2023 · 被引用 49 次
- LUQ: Long-text Uncertainty Quantification for LLMsCaiqi Zhang, Fangyu Liu, Marco Basaldella, Nigel CollierEMNLP 2024 · 被引用 11 次
相关 Paper
- Addressing Pitfalls in the Evaluation of Uncertainty Estimation Methods for Natural Language GenerationMykyta Ielanskyi, Kajetan Schweighofer, Lukas Aichberger, Sepp HochreiterICLR 2026 · 被引用 10 次
- Calibrating the Confidence of Large Language Models by Eliciting FidelityMozhi Zhang, Mianqiu Huang, Rundong Shi, Linsen Guo 等EMNLP 2024 · 被引用 5 次
- Inv-Entropy: A Fully Probabilistic Framework for Uncertainty Quantification in Language ModelsHaoyi Song, Ruihan Ji, Naichen Shi, Fan Lai 等NeurIPS 2025 · 被引用 6 次
- Kernel Language Entropy: Fine-grained Uncertainty Quantification for LLMs from Semantic SimilaritiesAlexander Nikitin, Jannik Kossen, Yarin Gal, Pekka MarttinenNeurIPS 2024 · 被引用 197 次
- Mapping from Meaning: Addressing the Miscalibration of Prompt-Sensitive Language ModelsKyle Cox, Jiawei Xu, Yikun Han, Rong Xu 等AAAI 2025 · 被引用 6 次
