Reward Model Learning vs. Direct Policy Optimization: A Comparative Analysis of Learning from Human Preferences
Andi Nika, Debmalya Mandal, Parameswaran Kamalaruban, Georgios Tzannetos, Goran Radanovic, Adish Singla
Abstract
In this paper, we take a step towards a deeper understanding of learning from human preferences by systematically comparing the paradigm of reinforcement learning from human feedback (RLHF) with the recently proposed paradigm of direct preference optimization (DPO). We focus our attention on the class of loglinear policy parametrization and linear reward functions. In order to compare the two paradigms, we first derive minimax statistical bounds on the suboptimality gap induced by both RLHF and DPO, assuming access to an oracle that exactly solves the optimization problems. We provide a detailed discussion on the relative comparison between the two paradigms, simultaneously taking into account the sample size, policy and reward class dimensions, and the regularization temperature. Moreover, we extend our analysis to the approximate optimization setting and derive exponentially decaying convergence rates for both RLHF and DPO. Next, we analyze the setting where the ground-truth reward is not realizable and find that, while RLHF incurs a constant additional error, DPO retains its asymptotically decaying gap by just tuning the temperature accordingly. Finally, we extend our comparison to the Markov decision process setting, where we generalize our results with exact optimization. To the best of our knowledge, we are the first to provide such a comparative analysis for RLHF and DPO.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9bcfcfd5-a876-48cc-aa67-3508d1a5f548Cited by top-tier papers10
- Robust LLM Alignment via Distributionally Robust Direct Preference OptimizationZaiyan Xu, Sushil Vemuri, Kishan Panaganti, Dileep Kalathil et al.NeurIPS 2025 · 18 citations
- Doubly Robust Alignment for Large Language ModelsErhan Xu, Kai Ye, Hongyi Zhou, Luhan Zhu et al.NeurIPS 2025 · 14 citations
- Can DPO Learn Diverse Human Values? A Theoretical Scaling LawShawn Im, Sharon LiNeurIPS 2025 · 8 citations
- Distributionally Robust Reinforcement Learning from Human FeedbackDebmalya Mandal, Paulius Sasnauskas, Goran RadanovicICML 2026
- Right Now, Wrong Then: Non-Stationary Direct Preference Optimization under Preference DriftSeongho Son, William Bankes, Sayak Ray Chowdhury, Brooks Paige et al.ICML 2025
Builds on20
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Scaling Laws for Reward Model OveroptimizationLeo Gao, John Schulman, Jacob HiltonICML 2023 · 963 citations
- An Optimistic Perspective on Offline Reinforcement LearningRishabh Agarwal, Dale Schuurmans, Mohammad NorouziICML 2020 · 568 citations
Related papers
- Understanding the Performance Gap in Preference Learning: A Dichotomy of RLHF and DPORuizhe Shi, Minhak Song, Runlong Zhou, Zihan Zhang et al.ICML 2026
- Conditional Equivalence of DPO and RLHF: Assumptions, Failure Modes, and Provable AlignmentYonggang Zhang, Zhiqin Yang, Wei Xue, Dong Fang et al.ICML 2026
- Zeroth-Order Policy Gradient for Reinforcement Learning from Human Feedback without Reward InferenceQining Zhang, Lei YingICLR 2025
- Explicit Preference Optimization: No Need for an Implicit Reward ModelXiangkun Hu, Lemin Kong, Tong He, David WipfICML 2025
- Policy-labeled Preference Learning: Is Preference Enough for RLHF?Taehyun Cho, Seokhun Ju, Seungyub Han, Dohyeong Kim et al.ICML 2025
