Contextualized Sequence Likelihood: Enhanced Confidence Scores for Natural Language Generation
Zhen Lin, Shubhendu Trivedi, Jimeng Sun
摘要
The advent of large language models (LLMs) has dramatically advanced the state-of-the-art in numerous natural language generation tasks.For LLMs to be applied reliably, it is essential to have an accurate measure of their confidence.Currently, the most commonly used confidence score function is the likelihood of the generated sequence, which, however, conflates semantic and syntactic components.For instance, in question-answering (QA) tasks, an awkward phrasing of the correct answer might result in a lower probability prediction.Additionally, different tokens should be weighted differently depending on the context.In this work, we propose enhancing the predicted sequence probability by assigning different weights to various tokens using attention values elicited from the base LLM.By employing a validation set, we can identify the relevant attention heads, thereby significantly improving the reliability of the vanilla sequence probability confidence measure.We refer to this new score as the Contextualized Sequence Likelihood (CSL).CSL is easy to implement, fast to compute, and offers considerable potential for further improvement with task-specific prompts.Across several QA datasets and a diverse array of LLMs, CSL has demonstrated significantly higher reliability than state-of-the-art baselines in predicting generation quality, as measured by the AUROC or AUARC.* Context *: [ $ context ] [ additional question -answer pairs ]
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Efficient semantic uncertainty quantification in language models via diversity-steered samplingJi Won Park, Kyunghyun ChoNeurIPS 2025 · 被引用 3 次
- User-side Model Consistency Monitoring for Open Source Large Language Models Inference ServicesQijun Miao, Zhixuan FangACL 2025 · 被引用 1 次
- Diagnosing Multi-step Reasoning Failures in Black-box LLMs via Stepwise Confidence AttributionXiaoou Liu, Tiejin Chen, Dengjia Zhang, Yaqing Wang 等ICML 2026 · 被引用 1 次
- On LLMs’ Internal Representation of Code CorrectnessFrancisco Ribeiro, Claudio Spiess, Premkumar Devanbu, Sarah NadiICSE 2026
- Enhancing Uncertainty Estimation in LLMs with Expectation of Aggregated Internal BeliefZeguan Xiao, Diyang Dou, Boya Xiong, Yun Chen 等AAAI 2026
它引用的顶会 Paper11
- Large Language Models Are Zero-Shot Time Series ForecastersNate Gruver, Marc Finzi, Shikai Qiu, Andrew Gordon WilsonNeurIPS 2023 · 被引用 898 次
- Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMsMiao Xiong, Zhiyuan Hu, Xinyang Lu, Yifei Li 等ICLR 2024 · 被引用 867 次
- Uncertainty Estimation in Autoregressive Structured PredictionAndrey Malinin, Mark J. F. GalesICLR 2021 · 被引用 439 次
- Talk like a Graph: Encoding Graphs for Large Language ModelsBahare Fatemi, Jonathan Halcrow, Bryan PerozziICLR 2024 · 被引用 194 次
- Conformal Language ModelingVictor Quach, Adam Fisch, Tal Schuster, Adam Yala 等ICLR 2024 · 被引用 132 次
相关 Paper
- MARS: Meaning-Aware Response Scoring for Uncertainty Estimation in Generative LLMsYavuz Faruk Bakman, Duygu Nur Yaldiz, Baturalp Buyukates, Chenyang Tao 等ACL 2024 · 被引用 6 次
- Efficient Self-Evaluation for Diffusion Language Models via Sequence RegenerationLinhao Zhong, Linyu Wu, Wen Wang, Yuling Xi 等ACL 2026
- Shifting Attention to Relevance: Towards the Predictive Uncertainty Quantification of Free-Form Large Language ModelsJinhao Duan, Hao Cheng, Shiqi Wang, Alex Zavalny 等ACL 2024 · 被引用 28 次
- QRelScore: Better Evaluating Generated Questions with Deeper Understanding of Context-aware RelevanceXiaoqiang Wang, Bang Liu, Siliang Tang, Lingfei WuEMNLP 2022 · 被引用 6 次
- Learning to Route LLMs with Confidence TokensYu-Neng Chuang, Prathusha Kameswara Sarma, Parikshit Gopalan, John Boccio 等ICML 2025
