Language Models with Conformal Factuality Guarantees
Christopher Mohri, Tatsunori Hashimoto
摘要
Guaranteeing the correctness and factuality of language model (LM) outputs is a major open problem. In this work, we propose conformal factuality, a framework that can ensure high probability correctness guarantees for LMs by connecting language modeling and conformal prediction. We observe that the correctness of an LM output is equivalent to an uncertainty quantification problem, where the uncertainty sets are defined as the entailment set of an LM's output. Using this connection, we show that conformal prediction in language models corresponds to a back-off algorithm that provides high probability correctness guarantees by progressively making LM outputs less specific (and expanding the associated uncertainty sets). This approach applies to any black-box LM and requires very few human-annotated samples. Evaluations of our approach on closed book QA (FActScore, NaturalQuestions) and reasoning tasks (MATH) show that our approach can provide 80-90% correctness guarantees while retaining the majority of the LM's original output.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper54
- Large language model validity via enhanced conformal prediction methodsJohn J. Cherian, Isaac Gibbs, Emmanuel J. CandèsNeurIPS 2024 · 被引用 120 次
- Conformal Alignment: Knowing When to Trust Foundation Models with GuaranteesYu Gui, Ying Jin, Zhimei RenNeurIPS 2024 · 被引用 63 次
- Length Optimization in Conformal PredictionShayan Kiyani, George J. Pappas, Hamed HassaniNeurIPS 2024 · 被引用 48 次
- Knowledge Boundary of Large Language Models: A SurveyMoxin Li, Yong Zhao, Wenxuan Zhang, Shuaiyi Li 等ACL 2025 · 被引用 33 次
- Graph-based Uncertainty Metrics for Long-form Language Model GenerationsMingjian Jiang, Yangjun Ruan, Prasanna Sattigeri, Salim Roukos 等NeurIPS 2024 · 被引用 25 次
它引用的顶会 Paper15
- Improving Factuality and Reasoning in Language Models through Multiagent DebateYilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum 等ICML 2024 · 被引用 1,562 次
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le 等ICLR 2023 · 被引用 681 次
- SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language ModelsPotsawee Manakul, Adian Liusie, Mark J. F. GalesEMNLP 2023 · 被引用 331 次
- Factuality Enhanced Language Models for Open-Ended Text GenerationNayeon Lee, Wei Ping, Peng Xu, Mostofa Patwary 等NeurIPS 2022 · 被引用 318 次
- Conformal Risk ControlAnastasios Nikolas Angelopoulos, Stephen Bates, Adam Fisch, Lihua Lei 等ICLR 2024 · 被引用 242 次
相关 Paper
- Conformal Language Model Reasoning with Coherent FactualityMaxon Rubin-Toles, Maya Gambhir, Keshav Ramji, Aaron Roth 等ICLR 2025
- Conformal Prediction Beyond the Seen: A Missing Mass Perspective for Uncertainty Quantification in Generative ModelsSima Noorani, Shayan Kiyani, George J. Pappas, Hamed HassaniNeurIPS 2025 · 被引用 10 次
- Conformal Linguistic Calibration: Trading-off between Factuality and SpecificityZhengping Jiang, Anqi Liu, Benjamin Van DurmeNeurIPS 2025 · 被引用 22 次
- Differentiable Conformal Training for LLM Reasoning FactualityNathan Hittesdorf, Marco Salzetta, Lu ChengICML 2026
- Multi-LLM Adaptive Conformal Inference for Reliable LLM ResponseKangjun Noh, Seongchan Lee, Ilmun Kim, Kyungwoo SongICLR 2026
