SConU: Selective Conformal Uncertainty in Large Language Models
Zhiyuan Wang, Qingni Wang, Yue Zhang, Tianlong Chen, Xiaofeng Zhu, Xiaoshuang Shi, Kaidi Xu
摘要
As large language models are increasingly utilized in real-world applications, guarantees of task-specific metrics are essential for their reliable deployment. Previous studies have introduced various criteria of conformal uncertainty grounded in split conformal prediction, which offer user-specified correctness coverage. However, existing frameworks often fail to identify uncertainty data outliers that violate the exchangeability assumption, leading to unbounded miscoverage rates and unactionable prediction sets. In this paper, we propose a novel approach termed Selective Conformal Uncertainty (SConU), which, for the first time, implements significance tests, by developing two conformal p-values that are instrumental in determining whether a given sample deviates from the uncertainty distribution of the calibration set at a specific manageable risk level. Our approach not only facilitates rigorous management of miscoverage rates across both singledomain and interdisciplinary contexts, but also enhances the efficiency of predictions. Furthermore, we comprehensively analyze the components of the conformal procedures, aiming to approximate conditional coverage, particularly in high-stakes question-answering tasks. 1 * Corresponding Authors † Equal Contribution 1 The code implementation for our experiments is available at https://github.com/Zhiyuan-GG/SConU (a) Single-domain Miscalibration. bu sin es s la w ps yc ho lo gy bi ol og y ch em ist ry hi st or y ot he r he al th ec on om ic s m at h ph ys ic s co m pu te r sc ie nc e ph ilo so ph y en gi ne er in g bu si ne ss la w ps yc ho lo gy bi ol og y ch em is tr y hi st or y ot he r he al th ec on om ic s m at h ph ys ic s co m pu te r sc ie nc e ph ilo so ph y en gi ne er in g 0.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- SAFER: Risk-Constrained Sample-then-Filter in Large Language ModelsQingni Wang, Yue Fan, Xin WangICLR 2026 · 被引用 8 次
- MM-Snowball: Evaluating and Mitigating Hallucination Snowballing in Multimodal Multi-Turn DialogueYue Jiang, Xue JIANG, Lihua Zhang, Zhiqiang Wang 等ICML 2026 · 被引用 1 次
- UNCLE: Benchmarking Uncertainty Expressions in Long-Form GenerationRuihan Yang, Caiqi Zhang, Zhisong Zhang, Xinting Huang 等EMNLP 2025
- LLaVA Steering: Visual Instruction Tuning with 500x Fewer Parameters through Modality Linear Representation-SteeringJinhe Bi, Yujun Wang, Haokun Chen, Xun Xiao 等ACL 2025
- COIN: Uncertainty-Guarding Selective Question Answering for Foundation Models with Provable Risk GuaranteesZhiyuan Wang, Jinhao Duan, Qingni Wang, Xiaofeng Zhu 等AAAI 2026
它引用的顶会 Paper15
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Class-Conditional Conformal Prediction with Many ClassesTiffany Ding, Anastasios Angelopoulos, Stephen Bates, Michael I. Jordan 等NeurIPS 2023 · 被引用 160 次
- Conformal Language ModelingVictor Quach, Adam Fisch, Tal Schuster, Adam Yala 等ICLR 2024 · 被引用 132 次
- Large language model validity via enhanced conformal prediction methodsJohn J. Cherian, Isaac Gibbs, Emmanuel J. CandèsNeurIPS 2024 · 被引用 120 次
- Language Models with Conformal Factuality GuaranteesChristopher Mohri, Tatsunori HashimotoICML 2024 · 被引用 107 次
相关 Paper
- Domain-Shift-Aware Conformal Prediction for Large Language ModelsZhexiao Lin, Yuanyuan Li, Neeraj Sarna, Yuanyuan Gao 等ICML 2026 · 被引用 6 次
- Conformal Prediction Beyond the Seen: A Missing Mass Perspective for Uncertainty Quantification in Generative ModelsSima Noorani, Shayan Kiyani, George J. Pappas, Hamed HassaniNeurIPS 2025 · 被引用 10 次
- Prune 'n Predict: Optimizing LLM Decision-making with Conformal PredictionHarit Vishwakarma, Alan Mishler, Thomas Cook, Niccolò Dalmasso 等ICML 2025
- Multivariate Conformal SelectionTian Bai, Yue Zhao, Xiang Yu, Archer Y. YangICML 2025
- Quantifying and Understanding Uncertainty in Large Reasoning ModelsYangyi Li, Chenxu Zhao, Mengdi HuaiACL 2026
