Reward Model Learning vs. Direct Policy Optimization: A Comparative Analysis of Learning from Human Preferences
Andi Nika, Debmalya Mandal, Parameswaran Kamalaruban, Georgios Tzannetos, Goran Radanovic, Adish Singla
摘要
In this paper, we take a step towards a deeper understanding of learning from human preferences by systematically comparing the paradigm of reinforcement learning from human feedback (RLHF) with the recently proposed paradigm of direct preference optimization (DPO). We focus our attention on the class of loglinear policy parametrization and linear reward functions. In order to compare the two paradigms, we first derive minimax statistical bounds on the suboptimality gap induced by both RLHF and DPO, assuming access to an oracle that exactly solves the optimization problems. We provide a detailed discussion on the relative comparison between the two paradigms, simultaneously taking into account the sample size, policy and reward class dimensions, and the regularization temperature. Moreover, we extend our analysis to the approximate optimization setting and derive exponentially decaying convergence rates for both RLHF and DPO. Next, we analyze the setting where the ground-truth reward is not realizable and find that, while RLHF incurs a constant additional error, DPO retains its asymptotically decaying gap by just tuning the temperature accordingly. Finally, we extend our comparison to the Markov decision process setting, where we generalize our results with exact optimization. To the best of our knowledge, we are the first to provide such a comparative analysis for RLHF and DPO.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Robust LLM Alignment via Distributionally Robust Direct Preference OptimizationZaiyan Xu, Sushil Vemuri, Kishan Panaganti, Dileep Kalathil 等NeurIPS 2025 · 被引用 18 次
- Doubly Robust Alignment for Large Language ModelsErhan Xu, Kai Ye, Hongyi Zhou, Luhan Zhu 等NeurIPS 2025 · 被引用 14 次
- Can DPO Learn Diverse Human Values? A Theoretical Scaling LawShawn Im, Sharon LiNeurIPS 2025 · 被引用 8 次
- Distributionally Robust Reinforcement Learning from Human FeedbackDebmalya Mandal, Paulius Sasnauskas, Goran RadanovicICML 2026
- Right Now, Wrong Then: Non-Stationary Direct Preference Optimization under Preference DriftSeongho Son, William Bankes, Sayak Ray Chowdhury, Brooks Paige 等ICML 2025
它引用的顶会 Paper20
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- Scaling Laws for Reward Model OveroptimizationLeo Gao, John Schulman, Jacob HiltonICML 2023 · 被引用 963 次
- An Optimistic Perspective on Offline Reinforcement LearningRishabh Agarwal, Dale Schuurmans, Mohammad NorouziICML 2020 · 被引用 568 次
相关 Paper
- Understanding the Performance Gap in Preference Learning: A Dichotomy of RLHF and DPORuizhe Shi, Minhak Song, Runlong Zhou, Zihan Zhang 等ICML 2026
- Conditional Equivalence of DPO and RLHF: Assumptions, Failure Modes, and Provable AlignmentYonggang Zhang, Zhiqin Yang, Wei Xue, Dong Fang 等ICML 2026
- Zeroth-Order Policy Gradient for Reinforcement Learning from Human Feedback without Reward InferenceQining Zhang, Lei YingICLR 2025
- Explicit Preference Optimization: No Need for an Implicit Reward ModelXiangkun Hu, Lemin Kong, Tong He, David WipfICML 2025
- Policy-labeled Preference Learning: Is Preference Enough for RLHF?Taehyun Cho, Seokhun Ju, Seungyub Han, Dohyeong Kim 等ICML 2025
