Conformal Alignment: Knowing When to Trust Foundation Models with Guarantees
Yu Gui, Ying Jin, Zhimei Ren
摘要
Before deploying outputs from foundation models in high-stakes tasks, it is imperative to ensure that they align with human values. For instance, in radiology report generation, reports generated by a vision-language model must align with human evaluations before their use in medical decision-making. This paper presents Conformal Alignment, a general framework for identifying units whose outputs meet a user-specified alignment criterion. It is guaranteed that on average, a prescribed fraction of selected units indeed meet the alignment criterion, regardless of the foundation model or the data distribution. Given any pre-trained model and new units with model-generated outputs, Conformal Alignment leverages a set of reference data with ground-truth alignment status to train an alignment predictor. It then selects new units whose predicted alignment scores surpass a data-dependent threshold, certifying their corresponding outputs as trustworthy. Through applications to question answering and radiology report generation, we demonstrate that our method is able to accurately identify units with trustworthy outputs via lightweight training over a moderate amount of reference data. En route, we investigate the informativeness of various features in alignment prediction and combine them with standard models to construct the alignment predictor.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper24
- Conformal Linguistic Calibration: Trading-off between Factuality and SpecificityZhengping Jiang, Anqi Liu, Benjamin Van DurmeNeurIPS 2025 · 被引用 22 次
- Selective Generation for Controllable Language ModelsMinjae Lee, Kyungmin Kim, Taesoo Kim, Sangdon ParkNeurIPS 2024 · 被引用 21 次
- Conformalized Time Series with Semantic FeaturesBaiting Chen, Zhimei Ren, Lu ChengNeurIPS 2024 · 被引用 19 次
- SConU: Selective Conformal Uncertainty in Large Language ModelsZhiyuan Wang, Qingni Wang, Yue Zhang, Tianlong Chen 等ACL 2025 · 被引用 18 次
- Conditional Quantile Adjusted Conformal Prediction for Time SeriesCheng Yu, Zhoufan Zhu, Ke ZhuICML 2026 · 被引用 11 次
它引用的顶会 Paper11
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 被引用 3,729 次
- Consistent Estimators for Learning to Defer to an ExpertHussein Mozannar, David A. SontagICML 2020 · 被引用 267 次
- Conformal Language ModelingVictor Quach, Adam Fisch, Tal Schuster, Adam Yala 等ICLR 2024 · 被引用 132 次
相关 Paper
- Towards Statistical Factuality Guarantee for Large Vision-Language ModelsZhuohang Li, Chao Yan, Nicholas J. Jackson, Wendi Cui 等EMNLP 2025 · 被引用 2 次
- Conformal Arbitrage: Risk-Controlled Balancing of Competing Objectives in Language ModelsWilliam Overman, Mohsen BayatiNeurIPS 2025 · 被引用 12 次
- Conf-Gen: Conformal Uncertainty Quantification for Generative ModelsGabriel Loaiza-Ganem, Kevin Zhang, Wei Cui, Marc Law 等ICML 2026 · 被引用 1 次
- Language Models with Conformal Factuality GuaranteesChristopher Mohri, Tatsunori HashimotoICML 2024 · 被引用 107 次
- Utility-Directed Conformal Prediction: A Decision-Aware Framework for Actionable Uncertainty QuantificationSantiago Cortes-Gomez, Carlos Miguel Patiño, Yewon Byun, Steven Wu 等ICLR 2025
