Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences?
Paul Gölz, Nika Haghtalab, Kunhe Yang
Abstract
After pre-training, large language models are aligned with human preferences based on pairwise comparisons. State-of-the-art alignment methods (such as PPO-based RLHF and DPO) are built on the assumption of aligning with a single preference model, despite being deployed in settings where users have diverse preferences. As a result, it is not even clear that these alignment methods produce models that satisfy users on average -- a minimal requirement for pluralistic alignment. Drawing on social choice theory and modeling users' comparisons through individual Bradley-Terry (BT) models, we introduce an alignment method's distortion: the worst-case ratio between the optimal achievable average utility, and the average utility of the learned policy. The notion of distortion helps draw sharp distinctions between alignment methods: Nash Learning from Human Feedback achieves the minimax optimal distortion of (for the BT temperature ), robustly across utility distributions, distributions of comparison pairs, and permissible KL divergences from the reference policy. RLHF and DPO, by contrast, suffer distortion already without a KL constraint, and or even unbounded distortion in the full setting, depending on how comparison pairs are sampled.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7beff5b2-3820-4436-ae44-7230975ba7aaCited by top-tier papers7
- Strategic Candidacy in Generative AI ArenasChris Hays, Rachel Li, Bailey Flanigan, Manish RaghavanICML 2026 · 3 citations
- The Sign Estimator: Preference Modeling for LLM Alignment under HeterogeneityAli Aouad, Aymane El Gadarri, Vivek FariasICML 2026 · 2 citations
- Enforcing Axioms for AI Alignment under Loss-Based RulesAlexandros Hollender, Sonja KraiczyICLR 2026
- Pluralistic LeaderboardsNika Haghtalab, Ariel Procaccia, Han Shao, Serena Wang et al.ICML 2026
- Mind the Gap: Structure-Aware Consistency in Preference LearningMehryar Mohri, Yutao ZhongICML 2026
Builds on21
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos et al.ICML 2024 · 1,212 citations
- SimPO: Simple Preference Optimization with a Reference-Free RewardYu Meng, Mengzhou Xia, Danqi ChenNeurIPS 2024 · 1,203 citations
- Model Alignment as Prospect Theoretic OptimizationKawin Ethayarajh, Winnie Xu, Niklas Muennighoff, Dan Jurafsky et al.ICML 2024 · 973 citations
Related papers
- Distortion of AI Alignment Revisited: RLHF is a Decent Utilitarian AlignerKazusato Oko, Annie Ulichney, Nika Haghtalab, Han BaoICML 2026
- Improving LLM General Preference Alignment via Optimistic Online Mirror DescentYuheng Zhang, Dian Yu, Tao Ge, Linfeng Song et al.NeurIPS 2025 · 27 citations
- Multiplayer Nash Preference OptimizationFang Wu, Xu Huang, Weihao Xuan, Zhiwei Zhang et al.ICLR 2026 · 8 citations
- Doubly Robust Alignment for Large Language ModelsErhan Xu, Kai Ye, Hongyi Zhou, Luhan Zhu et al.NeurIPS 2025 · 14 citations
- How RLHF Amplifies SycophancyItai Shapira, Gerdus Benade, Ariel ProcacciaICML 2026 · 16 citations
