Aligning Model Properties via Conformal Risk Control
William Overman, Jacqueline Jil Vallon, Mohsen Bayati
Abstract
AI model alignment is crucial due to inadvertent biases in training data and the underspecified machine learning pipeline, where models with excellent test metrics may not meet end-user requirements. While post-training alignment via human feedback shows promise, these methods are often limited to generative AI settings where humans can interpret and provide feedback on model outputs. In traditional non-generative settings with numerical or categorical outputs, detecting misalignment through single-sample outputs remains challenging, and enforcing alignment during training requires repeating costly training processes. In this paper we consider an alternative strategy. We propose interpreting model alignment through property testing, defining an aligned model as one belonging to a subset of functions that exhibit specific desired behaviors. We focus on post-processing a pre-trained model to better align with using conformal risk control. Specifically, we develop a general procedure for converting queries for testing a given property to a collection of loss functions suitable for use in a conformal risk control algorithm. We prove a probabilistic guarantee that the resulting conformal interval around contains a function approximately satisfying . We exhibit applications of our methodology on a collection of supervised learning datasets for (shape-constrained) properties such as monotonicity and concavity. The general procedure is flexible and can be applied to a wide range of desired properties. Finally, we prove that pre-trained models will always require alignment techniques even as model sizes or training data increase, as long as the training data contains even small biases.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 16badf0f-e08f-496f-b7b2-059bfcf76ee6Cited by top-tier papers4
- Conformal Arbitrage: Risk-Controlled Balancing of Competing Objectives in Language ModelsWilliam Overman, Mohsen BayatiNeurIPS 2025 · 12 citations
- Confident and Adaptive Generative Speech Recognition via Risk ControlAmit Damri, Bracha Laufer-GoldshteinICLR 2026
- Multimodal Learning on Low-Quality Data with Conformal Predictive Self-CalibrationXun Jiang, Yufan Gu, Disen Hu, Yuqing Hou et al.CVPR 2026
- Calibrating Conservatism for Scalable OversightWilliam Overman, Mohsen BayatiICML 2026
Builds on8
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Aligning AI With Shared Human ValuesDan Hendrycks, Collin Burns, Steven Basart, Andrew Critch et al.ICLR 2021 · 878 citations
- The Alignment Problem from a Deep Learning PerspectiveRichard Ngo, Lawrence Chan, Sören MindermannICLR 2024 · 296 citations
- Conformal Risk ControlAnastasios Nikolas Angelopoulos, Stephen Bates, Adam Fisch, Lihua Lei et al.ICLR 2024 · 242 citations
- Goal Misgeneralization in Deep Reinforcement LearningLauro Langosco di Langosco, Jack Koch, Lee D. Sharkey, Jacob Pfau et al.ICML 2022 · 128 citations
Related papers
- Conformal Policy ControlDrew Prinster, Clara Fannjiang, Ji Won Park, Kyunghyun Cho et al.ICML 2026 · 3 citations
- Conf-Gen: Conformal Uncertainty Quantification for Generative ModelsGabriel Loaiza-Ganem, Kevin Zhang, Wei Cui, Marc Law et al.ICML 2026 · 1 citation
- Conformal Risk Training: End-to-End Optimization of Conformal Risk ControlChristopher Yeh, Nicolas Christianson, Adam Wierman, Yisong YueNeurIPS 2025 · 14 citations
- Conformal Validity Guarantees Exist for Any Data Distribution (and How to Find Them)Drew Prinster, Samuel Don Stanton, Anqi Liu, Suchi SariaICML 2024 · 20 citations
- Boosted Conformal Prediction IntervalsRan Xie, Rina Barber, Emmanuel J. CandèsNeurIPS 2024 · 34 citations
