Large language model validity via enhanced conformal prediction methods
John J. Cherian, Isaac Gibbs, Emmanuel J. Candès
Abstract
We develop new conformal inference methods for obtaining validity guarantees on the output of large language models (LLMs). Prior work in conformal language modeling identifies a subset of the text that satisfies a high-probability guarantee of correctness. These methods work by filtering claims from the LLM's original response if a scoring function evaluated on the claim fails to exceed a threshold calibrated via split conformal prediction. Existing methods in this area suffer from two deficiencies. First, the guarantee stated is not conditionally valid. The trustworthiness of the filtering step may vary based on the topic of the response. Second, because the scoring function is imperfect, the filtering step can remove many valuable and accurate claims. We address both of these challenges via two new conformal methods. First, we generalize the conditional conformal procedure of Gibbs et al. (2023) in order to adaptively issue weaker guarantees when they are required to preserve the utility of the output. Second, we show how to systematically improve the quality of the scoring function via a novel algorithm for differentiating through the conditional conformal procedure. We demonstrate the efficacy of our approach on biography and medical question-answering datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7517b59c-a16a-4885-abc6-90f91ff74696Cited by top-tier papers44
- Length Optimization in Conformal PredictionShayan Kiyani, George J. Pappas, Hamed HassaniNeurIPS 2024 · 48 citations
- Knowledge Boundary of Large Language Models: A SurveyMoxin Li, Yong Zhao, Wenxuan Zhang, Shuaiyi Li et al.ACL 2025 · 33 citations
- Conformal Linguistic Calibration: Trading-off between Factuality and SpecificityZhengping Jiang, Anqi Liu, Benjamin Van DurmeNeurIPS 2025 · 22 citations
- Selective Generation for Controllable Language ModelsMinjae Lee, Kyungmin Kim, Taesoo Kim, Sangdon ParkNeurIPS 2024 · 21 citations
- Backward Conformal PredictionEtienne Gauthier, Francis Bach, Michael I. JordanNeurIPS 2025 · 19 citations
Builds on9
- SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language ModelsPotsawee Manakul, Adian Liusie, Mark J. F. GalesEMNLP 2023 · 331 citations
- Conformal Risk ControlAnastasios Nikolas Angelopoulos, Stephen Bates, Adam Fisch, Lihua Lei et al.ICLR 2024 · 242 citations
- FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text GenerationSewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis et al.EMNLP 2023 · 225 citations
- Conformal Language ModelingVictor Quach, Adam Fisch, Tal Schuster, Adam Yala et al.ICLR 2024 · 132 citations
- Learning Optimal Conformal ClassifiersDavid Stutz, Krishnamurthy Dvijotham, Ali Taylan Cemgil, Arnaud DoucetICLR 2022 · 123 citations
Related papers
- Multi-LLM Adaptive Conformal Inference for Reliable LLM ResponseKangjun Noh, Seongchan Lee, Ilmun Kim, Kyungwoo SongICLR 2026
- Language Models with Conformal Factuality GuaranteesChristopher Mohri, Tatsunori HashimotoICML 2024 · 107 citations
- Inference-Time Conformal Reasoning with Valid Factuality Control for Large Language ModelsTing Wang, Yuanjie Shi, Yan Yan, Huan ZhangICML 2026
- Differentiable Conformal Training for LLM Reasoning FactualityNathan Hittesdorf, Marco Salzetta, Lu ChengICML 2026
- Conformal Language Model Reasoning with Coherent FactualityMaxon Rubin-Toles, Maya Gambhir, Keshav Ramji, Aaron Roth et al.ICLR 2025
