The Axiomatic Value of Regularization in AI Alignment from Human Preferences
Ezgi Korkmaz
Abstract
Reinforcement learning from human feedback is the leading approach to aligning powerful AI systems so that they can be safe and helpful for humanity. While RLHF is typically modelled as a problem of learning a single preference ranking from noisy feedback, true human preferences are complex and often conflicting, representing substantive disagreements stemming from the diversity of individual human values. With this motivation, a recent line of research has studied RLHF from the perspective of social choice theory, which provides a set of well-established desirable properties for aggregating diverse preferences. Seen through this lens, the standard learning objective in RLHF is equivalent to aggregating diverse human preferences via the Borda count rule. At the same time, several new RLHF algorithms have been proposed, which turn out to be equivalent to the von Neumann winner social choice rule. However, the connection between social choice theory and RLHF has thus far ignored the critical role of regularization to prevent divergence from a reference policy, which is utilized in essentially all practical RLHF algorithms. In this paper, we study how regularization affects the social choice axioms satisfied by different RLHF algorithms, and prove that regularization improves the axiomatic properties of the von Neumann winner rule. In contrast, the Borda count rule still fails to satisfy key social choice axioms even when regularized. These results provide a principled argument grounded in social choice theory for utilizing practical RLHF algorithms that correspond to the von Neumann winner, rather than the standard RLHF objective.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on6
- Nash Learning from Human FeedbackRémi Munos, Michal Valko, Daniele Calandriello, Mohammad Gheshlaghi Azar et al.ICML 2024 · 212 citations
- A Minimaximalist Approach to Reinforcement Learning from Human FeedbackGokul Swamy, Christoph Dann, Rahul Kidambi, Steven Wu et al.ICML 2024 · 147 citations
- Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHFAnand Siththaranjan, Cassidy Laidlaw, Dylan Hadfield-MenellICLR 2024 · 112 citations
- Axioms for AI Alignment from Human FeedbackLuise Ge, Daniel Halpern, Evi Micha, Ariel D. Procaccia et al.NeurIPS 2024 · 64 citations
- Fair Reinforcement Learning for Just AIEzgi KorkmazICLR 2026
Related papers
- RA-PbRL: Provably Efficient Risk-Aware Preference-Based Reinforcement LearningYujie Zhao, Jose E. Aguilar Escamilla, Weyl Lu, Huazheng WangNeurIPS 2024 · 11 citations
- Is RLHF More Difficult than Standard RL? A Theoretical PerspectiveYuanhao Wang, Qinghua Liu, Chi JinNeurIPS 2023
- Strategyproof Reinforcement Learning from Human FeedbackThomas Kleine Buening, Jiarui Gan, Debmalya Mandal, Marta KwiatkowskaNeurIPS 2025 · 10 citations
- Contrastive Preference Learning: Learning from Human Feedback without Reinforcement LearningJoey Hejna, Rafael Rafailov, Harshit Sikchi, Chelsea Finn et al.ICLR 2024 · 37 citations
- MaxMin-RLHF: Alignment with Diverse Human PreferencesSouradip Chakraborty, Jiahao Qiu, Hui Yuan, Alec Koppel et al.ICML 2024 · 104 citations
