Multi-Reference Preference Optimization for Large Language Models
Hung Le, Quan Hung Tran, Dung Nguyen, Kien Do, Saloni Mittal, Kelechi Ogueji, Svetha Venkatesh
摘要
How can Large Language Models (LLMs) be aligned with human intentions and values? A typical solution is to gather human preference on model outputs and finetune the LLMs accordingly while ensuring that updates do not deviate too far from a reference model. Recent approaches, such as direct preference optimization (DPO), have eliminated the need for unstable and sluggish reinforcement learning optimization by introducing close-formed supervised losses. However, a significant limitation of the current approach is its design for a single reference model only, neglecting to leverage the collective power of numerous pretrained LLMs. To overcome this limitation, we introduce a novel closed-form formulation for direct preference optimization using multiple reference models. The resulting algorithm, Multi-Reference Preference Optimization (MRPO), leverages broader prior knowledge from diverse reference models, substantially enhancing preference learning capabilities compared to the single-reference DPO. Our experiments demonstrate that LLMs finetuned with MRPO generalize better in various preference data, regardless of data scarcity or abundance. Furthermore, MRPO effectively finetunes LLMs to exhibit superior performance in several downstream natural language processing tasks such as GSM8K and TruthfulQA. Recent advancements, such as direct preference optimization (DPO [Rafailov et al., 2023]) and other likelihood-based preference learning [
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- KL-Regularized RLHF with Multiple Reference Models: Exact Solutions and Sample ComplexityGholamali Aminian, Amir Reza Asadi, Idan Shenfeld, Youssef MrouehNeurIPS 2025 · 被引用 11 次
- Unifying Stable Optimization and Reference Regularization in RLHFLi He, Qiang Qu, He Zhao, Stephen Wan 等ICLR 2026 · 被引用 6 次
- The Realignment Problem: When Right becomes Wrong in LLMsAakash Sen Sharma, Debdeep Sanyal, Manodeep Ray, Vivek Srivastava 等ICML 2026
它引用的顶会 Paper9
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer 等NeurIPS 2022 · 被引用 2,039 次
- Model Alignment as Prospect Theoretic OptimizationKawin Ethayarajh, Winnie Xu, Niklas Muennighoff, Dan Jurafsky 等ICML 2024 · 被引用 973 次
- Self-Rewarding Language ModelsWeizhe Yuan, Richard Yuanzhe Pang, Kyunghyun Cho, Xian Li 等ICML 2024 · 被引用 569 次
相关 Paper
- Token-level Direct Preference OptimizationYongcheng Zeng, Guoqing Liu, Weiyu Ma, Ning Yang 等ICML 2024 · 被引用 136 次
- Expectation Preference Optimization: Reliable Preference Estimation for Improving the Reasoning Capability of Large Language ModelsZelin Li, Dawei SongEMNLP 2025
- Beyond Pairwise: Empowering LLM Alignment With (Ranked) Choice ModelingYuxuan Tang, Yifan FengICLR 2026 · 被引用 1 次
- Multiplayer Nash Preference OptimizationFang Wu, Xu Huang, Weihao Xuan, Zhiwei Zhang 等ICLR 2026 · 被引用 8 次
- Gradient-Adaptive Policy Optimization: Towards Multi-Objective Alignment of Large Language ModelsChengao Li, Hanyu Zhang, Yunkun Xu, Hongyan Xue 等ACL 2025 · 被引用 13 次
