Self-Improving Robust Preference Optimization
Eugene Choi, Arash Ahmadian, Matthieu Geist, Olivier Pietquin, Mohammad Gheshlaghi Azar
摘要
Online and offline RLHF methods, such as PPO and DPO, have been highly successful in aligning AI with human preferences. Despite their success, however, these methods suffer from fundamental limitations: (a) Models trained with RLHF can learn from mistakes or negative examples through RL mechanism or contrastive loss during training. However, at inference time, they lack an innate self-improvement mechanism for error corrections. (b) The optimal solution of existing methods is highly task-dependent, making it difficult for them to generalize to new tasks. To address these challenges, we propose Self-Improving Robust Preference Optimization (SRPO), a practical and mathematically principled offline RLHF framework. The key idea behind SRPO is to cast the problem of learning from human preferences as a self-improvement process, mathematically formulated as a min-max objective that jointly optimizes a self-improvement policy and a generative policy in an adversarial fashion. Crucially, the solution for this optimization problem is independent of the training task, which makes it robust to its changes. We then show that this objective can be reformulated as a non-adversarial offline loss, which can be efficiently optimized using standard supervised learning techniques at scale. To demonstrate SRPO's effectiveness, we evaluate it using AI Win-Rate (WR) against human (GOLD) completions. When tested on the XSum dataset, SRPO outperforms DPO by a margin of 15% after 5 self revisions, achieving an impressive 90% WR. Moreover, on the challenging Arena-Hard prompts, SRPO outperforms both DPO and IPO (by 4% without revision and 6% after a single revision), reaching a 56% WR against against Llama-3.1-8B-Instruct.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Back to Basics: Revisiting REINFORCE-Style Optimization for Learning from Human Feedback in LLMsArash Ahmadian, Chris Cremer, Matthias Gallé, Marzieh Fadaee 等ACL 2024 · 被引用 20 次
- Sherlock: Self-Correcting Reasoning in Vision-Language ModelsYi Ding, Ruqi ZhangNeurIPS 2025 · 被引用 14 次
- Self-Evolutionary Large Language Models Through Uncertainty-Enhanced Preference OptimizationJianing Wang, Yang Zhou, Xiaocheng Zhang, Mengjiao Bao 等AAAI 2025 · 被引用 8 次
- RLHF Can Speak Many Languages: Unlocking Multilingual Preference Optimization for LLMsJohn Dang, Arash Ahmadian, Kelly Marchisio, Julia Kreutzer 等EMNLP 2024 · 被引用 4 次
- Safety Alignment of LMs via Non-cooperative GamesAnselm Paulus, Ilia Kulikov, Brandon Amos, REMI MUNOS 等ICML 2026 · 被引用 4 次
它引用的顶会 Paper24
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
相关 Paper
- Online Preference Alignment for Language Models via Count-based ExplorationChenjia Bai, Yang Zhang, Shuang Qiu, Qiaosheng Zhang 等ICLR 2025
- AlphaDPO: Adaptive Reward Margin for Direct Preference OptimizationJunkang Wu, Xue Wang, Zhengyi Yang, Jiancan Wu 等ICML 2025
- RRHF: Rank Responses to Align Language Models with Human FeedbackHongyi Yuan, Zheng Yuan, Chuanqi Tan, Wei Wang 等NeurIPS 2023 · 被引用 515 次
- Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHFTengyang Xie, Dylan J. Foster, Akshay Krishnamurthy, Corby Rosset 等ICLR 2025
- Is DPO Superior to PPO for LLM Alignment? A Comprehensive StudyShusheng Xu, Wei Fu, Jiaxuan Gao, Wenjie Ye 等ICML 2024 · 被引用 274 次
