Aligning Model Properties via Conformal Risk Control
William Overman, Jacqueline Jil Vallon, Mohsen Bayati
摘要
AI model alignment is crucial due to inadvertent biases in training data and the underspecified machine learning pipeline, where models with excellent test metrics may not meet end-user requirements. While post-training alignment via human feedback shows promise, these methods are often limited to generative AI settings where humans can interpret and provide feedback on model outputs. In traditional non-generative settings with numerical or categorical outputs, detecting misalignment through single-sample outputs remains challenging, and enforcing alignment during training requires repeating costly training processes. In this paper we consider an alternative strategy. We propose interpreting model alignment through property testing, defining an aligned model as one belonging to a subset of functions that exhibit specific desired behaviors. We focus on post-processing a pre-trained model to better align with using conformal risk control. Specifically, we develop a general procedure for converting queries for testing a given property to a collection of loss functions suitable for use in a conformal risk control algorithm. We prove a probabilistic guarantee that the resulting conformal interval around contains a function approximately satisfying . We exhibit applications of our methodology on a collection of supervised learning datasets for (shape-constrained) properties such as monotonicity and concavity. The general procedure is flexible and can be applied to a wide range of desired properties. Finally, we prove that pre-trained models will always require alignment techniques even as model sizes or training data increase, as long as the training data contains even small biases.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Conformal Arbitrage: Risk-Controlled Balancing of Competing Objectives in Language ModelsWilliam Overman, Mohsen BayatiNeurIPS 2025 · 被引用 12 次
- Confident and Adaptive Generative Speech Recognition via Risk ControlAmit Damri, Bracha Laufer-GoldshteinICLR 2026
- Multimodal Learning on Low-Quality Data with Conformal Predictive Self-CalibrationXun Jiang, Yufan Gu, Disen Hu, Yuqing Hou 等CVPR 2026
- Calibrating Conservatism for Scalable OversightWilliam Overman, Mohsen BayatiICML 2026
它引用的顶会 Paper8
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Aligning AI With Shared Human ValuesDan Hendrycks, Collin Burns, Steven Basart, Andrew Critch 等ICLR 2021 · 被引用 878 次
- The Alignment Problem from a Deep Learning PerspectiveRichard Ngo, Lawrence Chan, Sören MindermannICLR 2024 · 被引用 296 次
- Conformal Risk ControlAnastasios Nikolas Angelopoulos, Stephen Bates, Adam Fisch, Lihua Lei 等ICLR 2024 · 被引用 242 次
- Goal Misgeneralization in Deep Reinforcement LearningLauro Langosco di Langosco, Jack Koch, Lee D. Sharkey, Jacob Pfau 等ICML 2022 · 被引用 128 次
相关 Paper
- Conformal Policy ControlDrew Prinster, Clara Fannjiang, Ji Won Park, Kyunghyun Cho 等ICML 2026 · 被引用 3 次
- Conf-Gen: Conformal Uncertainty Quantification for Generative ModelsGabriel Loaiza-Ganem, Kevin Zhang, Wei Cui, Marc Law 等ICML 2026 · 被引用 1 次
- Conformal Risk Training: End-to-End Optimization of Conformal Risk ControlChristopher Yeh, Nicolas Christianson, Adam Wierman, Yisong YueNeurIPS 2025 · 被引用 14 次
- Conformal Validity Guarantees Exist for Any Data Distribution (and How to Find Them)Drew Prinster, Samuel Don Stanton, Anqi Liu, Suchi SariaICML 2024 · 被引用 20 次
- Boosted Conformal Prediction IntervalsRan Xie, Rina Barber, Emmanuel J. CandèsNeurIPS 2024 · 被引用 34 次
