Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences?
Paul Gölz, Nika Haghtalab, Kunhe Yang
摘要
After pre-training, large language models are aligned with human preferences based on pairwise comparisons. State-of-the-art alignment methods (such as PPO-based RLHF and DPO) are built on the assumption of aligning with a single preference model, despite being deployed in settings where users have diverse preferences. As a result, it is not even clear that these alignment methods produce models that satisfy users on average -- a minimal requirement for pluralistic alignment. Drawing on social choice theory and modeling users' comparisons through individual Bradley-Terry (BT) models, we introduce an alignment method's distortion: the worst-case ratio between the optimal achievable average utility, and the average utility of the learned policy. The notion of distortion helps draw sharp distinctions between alignment methods: Nash Learning from Human Feedback achieves the minimax optimal distortion of (for the BT temperature ), robustly across utility distributions, distributions of comparison pairs, and permissible KL divergences from the reference policy. RLHF and DPO, by contrast, suffer distortion already without a KL constraint, and or even unbounded distortion in the full setting, depending on how comparison pairs are sampled.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Strategic Candidacy in Generative AI ArenasChris Hays, Rachel Li, Bailey Flanigan, Manish RaghavanICML 2026 · 被引用 3 次
- The Sign Estimator: Preference Modeling for LLM Alignment under HeterogeneityAli Aouad, Aymane El Gadarri, Vivek FariasICML 2026 · 被引用 2 次
- Enforcing Axioms for AI Alignment under Loss-Based RulesAlexandros Hollender, Sonja KraiczyICLR 2026
- Pluralistic LeaderboardsNika Haghtalab, Ariel Procaccia, Han Shao, Serena Wang 等ICML 2026
- Mind the Gap: Structure-Aware Consistency in Preference LearningMehryar Mohri, Yutao ZhongICML 2026
它引用的顶会 Paper21
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos 等ICML 2024 · 被引用 1,212 次
- SimPO: Simple Preference Optimization with a Reference-Free RewardYu Meng, Mengzhou Xia, Danqi ChenNeurIPS 2024 · 被引用 1,203 次
- Model Alignment as Prospect Theoretic OptimizationKawin Ethayarajh, Winnie Xu, Niklas Muennighoff, Dan Jurafsky 等ICML 2024 · 被引用 973 次
相关 Paper
- Distortion of AI Alignment Revisited: RLHF is a Decent Utilitarian AlignerKazusato Oko, Annie Ulichney, Nika Haghtalab, Han BaoICML 2026
- Improving LLM General Preference Alignment via Optimistic Online Mirror DescentYuheng Zhang, Dian Yu, Tao Ge, Linfeng Song 等NeurIPS 2025 · 被引用 27 次
- Multiplayer Nash Preference OptimizationFang Wu, Xu Huang, Weihao Xuan, Zhiwei Zhang 等ICLR 2026 · 被引用 8 次
- Doubly Robust Alignment for Large Language ModelsErhan Xu, Kai Ye, Hongyi Zhou, Luhan Zhu 等NeurIPS 2025 · 被引用 14 次
- How RLHF Amplifies SycophancyItai Shapira, Gerdus Benade, Ariel ProcacciaICML 2026 · 被引用 16 次
