Conformal Alignment: Knowing When to Trust Foundation Models with Guarantees
Yu Gui, Ying Jin, Zhimei Ren
Abstract
Before deploying outputs from foundation models in high-stakes tasks, it is imperative to ensure that they align with human values. For instance, in radiology report generation, reports generated by a vision-language model must align with human evaluations before their use in medical decision-making. This paper presents Conformal Alignment, a general framework for identifying units whose outputs meet a user-specified alignment criterion. It is guaranteed that on average, a prescribed fraction of selected units indeed meet the alignment criterion, regardless of the foundation model or the data distribution. Given any pre-trained model and new units with model-generated outputs, Conformal Alignment leverages a set of reference data with ground-truth alignment status to train an alignment predictor. It then selects new units whose predicted alignment scores surpass a data-dependent threshold, certifying their corresponding outputs as trustworthy. Through applications to question answering and radiology report generation, we demonstrate that our method is able to accurately identify units with trustworthy outputs via lightweight training over a moderate amount of reference data. En route, we investigate the informativeness of various features in alignment prediction and combine them with standard models to construct the alignment predictor.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 467b4015-0697-4d0d-a6d9-2db26beebee3Cited by top-tier papers24
- Conformal Linguistic Calibration: Trading-off between Factuality and SpecificityZhengping Jiang, Anqi Liu, Benjamin Van DurmeNeurIPS 2025 · 22 citations
- Selective Generation for Controllable Language ModelsMinjae Lee, Kyungmin Kim, Taesoo Kim, Sangdon ParkNeurIPS 2024 · 21 citations
- Conformalized Time Series with Semantic FeaturesBaiting Chen, Zhimei Ren, Lu ChengNeurIPS 2024 · 19 citations
- SConU: Selective Conformal Uncertainty in Large Language ModelsZhiyuan Wang, Qingni Wang, Yue Zhang, Tianlong Chen et al.ACL 2025 · 18 citations
- Conditional Quantile Adjusted Conformal Prediction for Time SeriesCheng Yu, Zhoufan Zhu, Ke ZhuICML 2026 · 11 citations
Builds on11
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 3,729 citations
- Consistent Estimators for Learning to Defer to an ExpertHussein Mozannar, David A. SontagICML 2020 · 267 citations
- Conformal Language ModelingVictor Quach, Adam Fisch, Tal Schuster, Adam Yala et al.ICLR 2024 · 132 citations
Related papers
- Towards Statistical Factuality Guarantee for Large Vision-Language ModelsZhuohang Li, Chao Yan, Nicholas J. Jackson, Wendi Cui et al.EMNLP 2025 · 2 citations
- Conformal Arbitrage: Risk-Controlled Balancing of Competing Objectives in Language ModelsWilliam Overman, Mohsen BayatiNeurIPS 2025 · 12 citations
- Conf-Gen: Conformal Uncertainty Quantification for Generative ModelsGabriel Loaiza-Ganem, Kevin Zhang, Wei Cui, Marc Law et al.ICML 2026 · 1 citation
- Language Models with Conformal Factuality GuaranteesChristopher Mohri, Tatsunori HashimotoICML 2024 · 107 citations
- Utility-Directed Conformal Prediction: A Decision-Aware Framework for Actionable Uncertainty QuantificationSantiago Cortes-Gomez, Carlos Miguel Patiño, Yewon Byun, Steven Wu et al.ICLR 2025
