Prune 'n Predict: Optimizing LLM Decision-making with Conformal Prediction
Harit Vishwakarma, Alan Mishler, Thomas Cook, Niccolò Dalmasso, Natraj Raman, Sumitra Ganesh
摘要
Large language models (LLMs) are empowering decision-making in several applications, including tool or API usage and answering multiple-choice questions (MCQs). However, incorrect outputs pose significant risks in high-stakes domains like healthcare and finance. To quantify LLM uncertainty and thereby mitigate these risks, recent works employ conformal prediction (CP), a model-and distribution-agnostic framework that uses LLM outputs to generate a prediction set containing the true answer with high probability. Leveraging CP, we propose conformal revision of questions (CROQ), which revises the question by narrowing down the available choices to those in the prediction set and asking the LLM the revised question. We expect LLMs to be more accurate on revised questions with fewer choices. Furthermore, we expect CROQ to be effective when the prediction sets from CP are small. Commonly used logit scores often lead to large sets, diminishing CROQ's effectiveness. To overcome this, we propose CP-OPT, an optimization framework to learn scores that minimize set sizes while maintaining coverage. Our extensive experiments on MMLU, ToolAlpaca, and TruthfulQA datasets with multiple LLMs show that CROQ improves accuracy over the standard inference, with more pronounced gains when paired with CP-OPT † .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Domain-Shift-Aware Conformal Prediction for Large Language ModelsZhexiao Lin, Yuanyuan Li, Neeraj Sarna, Yuanyuan Gao 等ICML 2026 · 被引用 6 次
- RACER: Risk-Aware Calibrated Efficient Routing for Large Language ModelsSai Hao, Hao Zeng, Hongxin Wei, Bingyi JingICML 2026 · 被引用 1 次
- Adaptively Grouped Contextual Bandits for Heterogeneous Human-AI Decision Making with Conformal Prediction SetsYanchen Wu, Bo LiICML 2026
- Easier to Judge than to Find: Predicting In-Context Learning Success for Demonstration SelectionHaochun Wang, Chaofen Yang, Jiatong Liu, Jingbo Wang 等ICML 2026
- CAOS: Conformal Aggregation of One-Shot PredictorsMaja WaldronICML 2026
它引用的顶会 Paper13
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 被引用 3,228 次
- Large Language Models Are Not Robust Multiple Choice SelectorsChujie Zheng, Hao Zhou, Fandong Meng, Jie Zhou 等ICLR 2024 · 被引用 424 次
- Conformal Language ModelingVictor Quach, Adam Fisch, Tal Schuster, Adam Yala 等ICLR 2024 · 被引用 132 次
- Learning Optimal Conformal ClassifiersDavid Stutz, Krishnamurthy Dvijotham, Ali Taylan Cemgil, Arnaud DoucetICLR 2022 · 被引用 123 次
相关 Paper
- Conformal Prediction Beyond the Seen: A Missing Mass Perspective for Uncertainty Quantification in Generative ModelsSima Noorani, Shayan Kiyani, George J. Pappas, Hamed HassaniNeurIPS 2025 · 被引用 10 次
- Language Models with Conformal Factuality GuaranteesChristopher Mohri, Tatsunori HashimotoICML 2024 · 被引用 107 次
- Conformal Constrained Policy Optimization for Cost-Effective LLM AgentsWenwen Si, Sooyong Jang, Insup Lee, Osbert BastaniAAAI 2026 · 被引用 3 次
- Analyzing Uncertainty of LLM-as-a-Judge: Interval Evaluations with Conformal PredictionHuanxin Sheng, Xinyi Liu, Hangfeng He, Jieyu Zhao 等EMNLP 2025 · 被引用 1 次
- Conf-Gen: Conformal Uncertainty Quantification for Generative ModelsGabriel Loaiza-Ganem, Kevin Zhang, Wei Cui, Marc Law 等ICML 2026 · 被引用 1 次
