Direct Alignment with Heterogeneous Preferences
Ali Shirali, Arash Nasr-Esfahany, Abdullah Omar Alomar, Parsa Mirtaheri, Rediet Abebe, Ariel D. Procaccia
Abstract
Alignment with human preferences is commonly framed using a universal reward function, even though human preferences are inherently heterogeneous. We formalize this heterogeneity by introducing user types and examine the limits of the homogeneity assumption. We show that aligning to heterogeneous preferences with a single policy is best achieved using the average reward across user types. However, this requires additional information about annotators. We examine improvements under different information settings, focusing on direct alignment methods. We find that minimal information can yield first-order improvements, while full feedback from each user type leads to consistent learning of the optimal policy. Surprisingly, however, no sample-efficient consistent direct loss exists in this latter setting. These results reveal a fundamental tension between consistency and sample efficiency in direct policy alignment.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext aa6af1dc-4bd7-41cb-9952-6b03e5579d4dCited by top-tier papers7
- Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences?Paul Gölz, Nika Haghtalab, Kunhe YangNeurIPS 2025 · 29 citations
- Pairwise Calibrated Rewards for Pluralistic AlignmentDaniel Halpern, Evi Micha, Ariel D. Procaccia, Itai ShapiraNeurIPS 2025 · 15 citations
- Can DPO Learn Diverse Human Values? A Theoretical Scaling LawShawn Im, Sharon LiNeurIPS 2025 · 8 citations
- Preference learning made easy: Everything should be understood through win rateLily H. Zhang, Rajesh RanganathICML 2025
- Pluralistic LeaderboardsNika Haghtalab, Ariel Procaccia, Han Shao, Serena Wang et al.ICML 2026
Builds on28
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- SimPO: Simple Preference Optimization with a Reference-Free RewardYu Meng, Mengzhou Xia, Danqi ChenNeurIPS 2024 · 1,203 citations
- Whose Opinions Do Language Models Reflect?Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee et al.ICML 2023 · 764 citations
Related papers
- Reward Learning from Multiple Feedback TypesYannick Metz, András Geiszl, Raphaël Baur, Mennatallah El-AssadyICLR 2025
- Personalizing Reinforcement Learning from Human Feedback with Variational Preference LearningSriyash Poddar, Yanming Wan, Hamish Ivison, Abhishek Gupta et al.NeurIPS 2024 · 188 citations
- Towards Efficient Exact Optimization of Language Model AlignmentHaozhe Ji, Cheng Lu, Yilin Niu, Pei Ke et al.ICML 2024 · 32 citations
- Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game PerspectiveHaichuan Wang, Tao Lin, Lingkai Kong, Ce Li et al.ICML 2026 · 3 citations
- Conditional Equivalence of DPO and RLHF: Assumptions, Failure Modes, and Provable AlignmentYonggang Zhang, Zhiqin Yang, Wei Xue, Dong Fang et al.ICML 2026
