Multiplayer Nash Preference Optimization
Fang Wu, Xu Huang, Weihao Xuan, Zhiwei Zhang, Yijia Xiao, Frank Wan, Xiaomin Li, Bing Hu, Peng Xia, Jure Leskovec, Yejin Choi
摘要
Reinforcement learning from human feedback (RLHF) has emerged as the standard paradigm for aligning large language models with human preferences. However, reward-based methods grounded in the Bradley–Terry assumption struggle to capture the nontransitivity and heterogeneity of real-world preferences. To address this, recent studies have reframed alignment as a two-player Nash game, giving rise to Nash learning from human feedback (NLHF). While this perspective has inspired algorithms such as INPO, ONPO, and EGPO that offer strong theoretical and empirical guarantees, they remain fundamentally restricted to two-player interactions, introducing a single-opponent bias that fails to capture the full complexity of realistic preference structures. This work introduces Multiplayer Nash Preference Optimization (MNPO), a novel framework that generalizes NLHF to the multiplayer regime. It formulates alignment as an -player game, where each policy competes against a population of opponents while being regularized toward a reference model. We demonstrate that MNPO inherits the equilibrium guarantees of two-player methods while enabling richer competitive dynamics and improved coverage of diverse preference structures. Comprehensive empirical evaluation shows that MNPO consistently outperforms existing NLHF baselines on instruction-following benchmarks, achieving superior alignment quality under heterogeneous annotator conditions and mixed-policy evaluation scenarios. Together, these results establish MNPO as a principled and scalable framework for aligning LLMs with complex, non-transitive human preferences. Code is available at https://github.com/smiles724/MNPO.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Asymptotic Universal Alignment: A New Alignment Framework via Test-Time ScalingYang Cai, Weiqiang ZhengICML 2026
- Transitivity Meets Cyclicity: Explicit Preference Decomposition for Dynamic Large Language Model AlignmentYucong Huang, Xiucheng Li, Kaiqi Zhao, Jing LiICML 2026
它引用的顶会 Paper28
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 被引用 3,037 次
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan 等NeurIPS 2025 · 被引用 2,828 次
相关 Paper
- MPO: An Efficient Post-Processing Framework for Mixing Diverse Preference AlignmentTianze Wang, Dongnan Gui, Yifan Hu, Shuhang Lin 等ICML 2025
- Iterative Nash Policy Optimization: Aligning LLMs with General Preferences via No-Regret LearningYuheng Zhang, Dian Yu, Baolin Peng, Linfeng Song 等ICLR 2025
- Magnetic Preference Optimization: Achieving Last-iterate Convergence for Language Model AlignmentMingzhi Wang, Chengdong Ma, Qizhi Chen, Linjian Meng 等ICLR 2025
- Improving LLM General Preference Alignment via Optimistic Online Mirror DescentYuheng Zhang, Dian Yu, Tao Ge, Linfeng Song 等NeurIPS 2025 · 被引用 27 次
- Reliability-Aware LLM Alignment from Inconsistent Human FeedbackJingyi Huang, Ruohan Zong, Yujun Feng, Liran Ma 等ICML 2026
