Conformal Language Modeling
Victor Quach, Adam Fisch, Tal Schuster, Adam Yala, Jae Ho Sohn, Tommi S. Jaakkola, Regina Barzilay
Abstract
We propose a novel approach to conformal prediction for generative language models (LMs). Standard conformal prediction produces prediction sets -- in place of single predictions -- that have rigorous, statistical performance guarantees. LM responses are typically sampled from the model's predicted distribution over the large, combinatorial output space of natural language. Translating this process to conformal prediction, we calibrate a stopping rule for sampling different outputs from the LM that get added to a growing set of candidates until we are confident that the output set is sufficient. Since some samples may be low-quality, we also simultaneously calibrate and apply a rejection rule for removing candidates from the output set to reduce noise. Similar to conformal prediction, we prove that the sampled set returned by our procedure contains at least one acceptable answer with high probability, while still being empirically precise (i.e., small) on average. Furthermore, within this set of candidate responses, we show that we can also accurately identify subsets of individual components -- such as phrases or sentences -- that are each independently correct (e.g., that are not"hallucinations"), again with statistical guarantees. We demonstrate the promise of our approach on multiple tasks in open-domain question answering, text summarization, and radiology report generation using different LM variants.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 45f7e602-d475-4409-b38a-590192118e85Cited by top-tier papers74
- Kernel Language Entropy: Fine-grained Uncertainty Quantification for LLMs from Semantic SimilaritiesAlexander Nikitin, Jannik Kossen, Yarin Gal, Pekka MarttinenNeurIPS 2024 · 197 citations
- Large language model validity via enhanced conformal prediction methodsJohn J. Cherian, Isaac Gibbs, Emmanuel J. CandèsNeurIPS 2024 · 120 citations
- Language Models with Conformal Factuality GuaranteesChristopher Mohri, Tatsunori HashimotoICML 2024 · 107 citations
- Conformal Alignment: Knowing When to Trust Foundation Models with GuaranteesYu Gui, Ying Jin, Zhimei RenNeurIPS 2024 · 63 citations
- Conformal Prediction for Deep Classifier via Label RankingJianguo Huang, Huajun Xi, Linjun Zhang, Huaxiu Yao et al.ICML 2024 · 50 citations
Builds on17
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 3,228 citations
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le et al.ICLR 2023 · 681 citations
- Confident Adaptive Language ModelingTal Schuster, Adam Fisch, Jai Gupta, Mostafa Dehghani et al.NeurIPS 2022 · 394 citations
Related papers
- Towards Statistical Factuality Guarantee for Large Vision-Language ModelsZhuohang Li, Chao Yan, Nicholas J. Jackson, Wendi Cui et al.EMNLP 2025 · 2 citations
- Conformal Prediction Sets with Limited False PositivesAdam Fisch, Tal Schuster, Tommi S. Jaakkola, Regina BarzilayICML 2022 · 33 citations
- Conformal Prediction Beyond the Seen: A Missing Mass Perspective for Uncertainty Quantification in Generative ModelsSima Noorani, Shayan Kiyani, George J. Pappas, Hamed HassaniNeurIPS 2025 · 10 citations
- Domain-Shift-Aware Conformal Prediction for Large Language ModelsZhexiao Lin, Yuanyuan Li, Neeraj Sarna, Yuanyuan Gao et al.ICML 2026 · 6 citations
- Conf-Gen: Conformal Uncertainty Quantification for Generative ModelsGabriel Loaiza-Ganem, Kevin Zhang, Wei Cui, Marc Law et al.ICML 2026 · 1 citation
