Pairwise Calibrated Rewards for Pluralistic Alignment
Daniel Halpern, Evi Micha, Ariel D. Procaccia, Itai Shapira
Abstract
Current alignment pipelines presume a single, universal notion of desirable behavior. However, human preferences often diverge across users, contexts, and cultures. As a result, disagreement collapses into the majority signal and minority perspectives are discounted. To address this, we propose reflecting diverse human preferences through a distribution over multiple reward functions, each inducing a distinct aligned policy. The distribution is learned directly from pairwise preference without annotator identifiers or predefined groups. Instead, annotator disagreements are treated as informative soft labels. Our central criterion is pairwise calibration: for every pair of candidate responses, the proportion of reward functions preferring one response matches the fraction of annotators with that preference. We prove that even a small outlier-free ensemble can accurately represent diverse preference distributions. Empirically, we introduce and validate a practical training heuristic to learn such ensembles, and demonstrate its effectiveness through improved calibration, implying a more faithful representation of pluralistic values.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8dd3e0fa-a04d-4921-b44c-8220133f2e65Cited by top-tier papers6
- How RLHF Amplifies SycophancyItai Shapira, Gerdus Benade, Ariel ProcacciaICML 2026 · 16 citations
- Beyond RLHF and NLHF: Population-Proportional Alignment under an Axiomatic FrameworkKihyun Kim, Jiawei Zhang, Asuman Ozdaglar, Pablo A. ParriloICLR 2026 · 5 citations
- Robust AI Evaluation through Maximal LotteriesHadi Khalaf, Serena Wang, Daniel Halpern, Itai Shapira et al.ICML 2026 · 2 citations
- Calibrated Preference Learning: The Case of Label RankingSanto Thies, Viktor Bengs, Timo Kaufmann, Sebastian Vollmer et al.ICML 2026
- When Distance Distracts: Representation Distance Bias in BT-Loss for Reward ModelsTong Xie, Ching-Yuan Bai, Yuanhao Ban, Yunqi Hong et al.ICML 2026
Builds on28
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Whose Opinions Do Language Models Reflect?Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee et al.ICML 2023 · 764 citations
- Fine-tuning language models to find agreement among humans with diverse preferencesMichiel A. Bakker, Martin J. Chadwick, Hannah Sheahan, Michael Henry Tessler et al.NeurIPS 2022 · 349 citations
- Understanding the Effects of RLHF on LLM Generalisation and DiversityRobert Kirk, Ishita Mediratta, Christoforos Nalmpantis, Jelena Luketina et al.ICLR 2024 · 332 citations
- Rewarded soups: towards Pareto-optimal alignment by interpolating weights fine-tuned on diverse rewardsAlexandre Ramé, Guillaume Couairon, Corentin Dancette, Jean-Baptiste Gaya et al.NeurIPS 2023 · 295 citations
Related papers
- Direct Alignment with Heterogeneous PreferencesAli Shirali, Arash Nasr-Esfahany, Abdullah Omar Alomar, Parsa Mirtaheri et al.NeurIPS 2025 · 26 citations
- Geometric-Averaged Preference Optimization for Soft Preference LabelsHiroki Furuta, Kuang-Huei Lee, Shixiang Shane Gu, Yutaka Matsuo et al.NeurIPS 2024 · 24 citations
- Diverse Preference Learning for Capabilities and AlignmentStewart Slocum, Asher Parker-Sartori, Dylan Hadfield-MenellICLR 2025
- No Preference Left Behind: Group Distributional Preference OptimizationBinwei Yao, Zefan Cai, Yun-Shiuan Chuang, Shanglin Yang et al.ICLR 2025
- Personalizing Reinforcement Learning from Human Feedback with Variational Preference LearningSriyash Poddar, Yanming Wan, Hamish Ivison, Abhishek Gupta et al.NeurIPS 2024 · 188 citations
