LUQ: Long-text Uncertainty Quantification for LLMs
Caiqi Zhang, Fangyu Liu, Marco Basaldella, Nigel Collier
摘要
Large Language Models (LLMs) have demonstrated remarkable capability in a variety of NLP tasks. However, LLMs are also prone to generate nonfactual content. Uncertainty Quantification (UQ) is pivotal in enhancing our understanding of a model's confidence on its generation, thereby aiding in the mitigation of nonfactual outputs. Existing research on UQ predominantly targets short text generation, typically yielding brief, word-limited responses. However, real-world applications frequently necessitate much longer responses. Our study first highlights the limitations of current UQ methods in handling long text generation. We then introduce LUQ with its two variations: LUQ-ATOMIC and LUQ-PAIR, a series of novel sampling-based UQ approaches specifically designed for long text. Our findings reveal that LUQ outperforms existing baseline methods in correlating with the model's factuality scores (negative coefficient of -0.85 observed for Gemini Pro). To further improve the factuality of LLM responses, we propose LUQ-ENSEMBLE, a method that ensembles responses from multiple models and selects the response with the lowest uncertainty. The ensembling method greatly improves the response factuality upon the best standalone LLM. 1 * Now at Google DeepMind. † Work done outside of Amazon.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- LoVeC: Reinforcement Learning for Better Verbalized Confidence in Long-Form GenerationCaiqi Zhang, Xiaochen Zhu, Chengzu Li, Nigel Collier 等ACL 2026 · 被引用 16 次
- Efficient Hallucination Detection for LLMs Using Uncertainty-Aware Attention HeadsArtem Vazhentsev, Lyudmila Rvanova, Gleb Kuzmin, Ekaterina Fadeeva 等ICML 2026 · 被引用 16 次
- Conformity in Large Language ModelsXiaochen Zhu, Caiqi Zhang, Tom Stafford, Nigel Collier 等ACL 2025 · 被引用 16 次
- Calibrating Verbalized Confidence with Self-Generated DistractorsVictor Wang, Elias Stengel-EskinICLR 2026 · 被引用 15 次
- Uncertainty in Language Models: Assessment through Rank-CalibrationXinmeng Huang, Shuo Li, Mengxin Yu, Matteo Sesia 等EMNLP 2024 · 被引用 9 次
它引用的顶会 Paper12
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMsMiao Xiong, Zhiyuan Hu, Xinyang Lu, Yifei Li 等ICLR 2024 · 被引用 867 次
- Uncertainty Estimation in Autoregressive Structured PredictionAndrey Malinin, Mark J. F. GalesICLR 2021 · 被引用 439 次
- DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding SharingPengcheng He, Jianfeng Gao, Weizhu ChenICLR 2023 · 被引用 394 次
- SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language ModelsPotsawee Manakul, Adian Liusie, Mark J. F. GalesEMNLP 2023 · 被引用 331 次
相关 Paper
- IUQ: Interrogative Uncertainty Quantification for Long-Form Large Language Model GenerationHaozhi Fan, Jinhao Duan, Kaidi XuACL 2026 · 被引用 1 次
- UNCERTAINTY-LINE: Length-Invariant Estimation of Uncertainty for Large Language ModelsRoman Vashurin, Maiya Goloburda, Preslav Nakov, Maxim PanovEMNLP 2025 · 被引用 1 次
- Semantic Uncertainty Quantification of Hallucinations in LLMs: A Quantum Tensor Network Based MethodPragatheeswaran Vipulanandan, Kamal Premaratne, Dilip SarkarICLR 2026 · 被引用 4 次
- Semantic Density: Uncertainty Quantification for Large Language Models through Confidence Measurement in Semantic SpaceXin Qiu, Risto MiikkulainenNeurIPS 2024
- Sampling-Free Uncertainty Quantification via Hidden State Dynamics in Language ModelsYixin Bu, Guanyun Zou, Renzhi Wang, Runze Xia 等AAAI 2026 · 被引用 2 次
