Generalized Preference Optimization: A Unified Approach to Offline Alignment
Yunhao Tang, Zhaohan Daniel Guo, Zeyu Zheng, Daniele Calandriello, Rémi Munos, Mark Rowland, Pierre Harvey Richemond, Michal Valko, Bernardo Ávila Pires, Bilal Piot
摘要
Offline preference optimization allows fine-tuning large models directly from offline data, and has proved effective in recent alignment practices. We propose generalized preference optimization (GPO), a family of offline losses parameterized by a general class of convex functions. GPO enables a unified view over preference optimization, encompassing existing algorithms such as DPO, IPO and SLiC as special cases, while naturally introducing new variants. The GPO framework also sheds light on how offline algorithms enforce regularization, through the design of the convex function that defines the loss. Our analysis and experiments reveal the connections and subtle differences between the offline regularization and the KL divergence regularization intended by the canonical RLHF formulation. In a controlled setting akin to Gao et al. ( 2023 ), we also show that different GPO variants achieve similar trade-offs between regularization and performance, though the optimal values of hyper-parameter might differ as predicted by theory. In all, our results present new algorithmic toolkits and empirical insights to alignment practitioners.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper60
- Normalized Rewards for Preference OptimizationShawn Im, Federico Danieli, Skyler Seto, Barry-John Theobald 等ICML 2026 · 被引用 571 次
- Scaling Laws for Reward Model Overoptimization in Direct Alignment AlgorithmsRafael Rafailov, Yaswanth Chittepu, Ryan Park, Harshit Sikchi 等NeurIPS 2024 · 被引用 169 次
- On Softmax Direct Preference Optimization for RecommendationYuxin Chen, Junfei Tan, An Zhang, Zhengyi Yang 等NeurIPS 2024 · 被引用 126 次
- Group Robust Preference Optimization in Reward-free RLHFShyam Sundhar Ramesh, Yifan Hu, Iason Chaimalas, Viraj Mehta 等NeurIPS 2024 · 被引用 122 次
- Human Alignment of Large Language Models through Online Preference OptimisationDaniele Calandriello, Zhaohan Daniel Guo, Rémi Munos, Mark Rowland 等ICML 2024 · 被引用 90 次
它引用的顶会 Paper8
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Scaling Laws for Reward Model OveroptimizationLeo Gao, John Schulman, Jacob HiltonICML 2023 · 被引用 963 次
- A Minimaximalist Approach to Reinforcement Learning from Human FeedbackGokul Swamy, Christoph Dann, Rahul Kidambi, Steven Wu 等ICML 2024 · 被引用 147 次
- Learning to summarize with human feedbackNisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel M. Ziegler 等NeurIPS 2020 · 被引用 124 次
相关 Paper
- Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference OptimizationAudrey Huang, Wenhao Zhan, Tengyang Xie, Jason D. Lee 等ICLR 2025
- Design Considerations in Offline Preference-based RLAlekh Agarwal, Christoph Dann, Teodor Vanislavov MarinovICML 2025
- Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-constraintWei Xiong, Hanze Dong, Chenlu Ye, Ziqi Wang 等ICML 2024 · 被引用 346 次
- Provably Mitigating Overoptimization in RLHF: Your SFT Loss is Implicitly an Adversarial RegularizerZhihan Liu, Miao Lu, Shenao Zhang, Boyi Liu 等NeurIPS 2024 · 被引用 119 次
- The Importance of Online Data: Understanding Preference Fine-tuning via CoverageYuda Song, Gokul Swamy, Aarti Singh, J. Andrew Bagnell 等NeurIPS 2024 · 被引用 63 次
