Domain-Shift-Aware Conformal Prediction for Large Language Models
Zhexiao Lin, Yuanyuan Li, Neeraj Sarna, Yuanyuan Gao, Michael von Gablenz
摘要
Large language models have achieved impressive performance across diverse tasks. However, their tendency to produce overconfident and factually incorrect outputs, known as hallucinations, poses risks in real world applications. Conformal prediction provides finite-sample, distribution-free coverage guarantees, but standard conformal prediction breaks down under domain shift, often leading to under-coverage and unreliable prediction sets. We propose a new framework called Domain-Shift-Aware Conformal Prediction (DS-CP). Our framework adapts conformal prediction to large language models under domain shift, by systematically reweighting calibration samples based on their proximity to the test prompt, thereby preserving validity while enhancing adaptivity. Our theoretical analysis and experiments on the MMLU benchmark demonstrate that the proposed method delivers more reliable coverage than standard conformal prediction, especially under substantial distribution shifts, while maintaining efficiency. This provides a practical step toward trustworthy uncertainty quantification for large language models in real-world deployment.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Beyond Reactivity: Proactive Adaptive Conformal Inference for Online LLM FactualityXinyu Liu, Jun WuICML 2026
- CAOS: Conformal Aggregation of One-Shot PredictorsMaja WaldronICML 2026
它引用的顶会 Paper13
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive SummarizationJingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. LiuICML 2020 · 被引用 2,453 次
- Improving Language Models by Retrieving from Trillions of TokensSebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai 等ICML 2022 · 被引用 1,629 次
- In Search of Lost Domain GeneralizationIshaan Gulrajani, David Lopez-PazICLR 2021 · 被引用 1,416 次
相关 Paper
- CoFact: Conformal Factuality Guarantees for Language Models under Covariate ShiftZirui Hu, Zheng Zhang, Yingjie Wang, Leszek Rutkowski 等ICLR 2026
- Towards Statistical Factuality Guarantee for Large Vision-Language ModelsZhuohang Li, Chao Yan, Nicholas J. Jackson, Wendi Cui 等EMNLP 2025 · 被引用 2 次
- Conditional Factuality Controlled LLMs with Generalization Certificates via Conformal SamplingKai Ye, Qingtao Pan, Shuo LiCVPR 2026 · 被引用 1 次
- Prune 'n Predict: Optimizing LLM Decision-making with Conformal PredictionHarit Vishwakarma, Alan Mishler, Thomas Cook, Niccolò Dalmasso 等ICML 2025
- Ensemble Conformal Predictor (EnCP): A New Conformal Predictor with Robustness Guarantees Against Data Poisoning AttacksYuxin Yang, Qiang Li, Runyang Feng, Liren Shan 等S&P 2026
