Learning Guarantee of Reward Modeling Using Deep Neural Networks
Yuanhang Luo, Yeheng Ge, Ruijian Han, Guohao Shen
Abstract
In this work, we study the learning theory of reward modeling with pairwise comparison data using deep neural networks. We establish a novel non-asymptotic regret bound for deep reward estimators in a non-parametric setting, which depends explicitly on the network architecture. Furthermore, to underscore the critical importance of clear human beliefs, we introduce a margin-type condition that assumes the conditional winning probability of the optimal action in pairwise comparisons is significantly distanced from 1/2. This condition enables a sharper regret bound, which substantiates the empirical efficiency of Reinforcement Learning from Human Feedback and highlights clear human beliefs in its success. Notably, this improvement stems from high-quality pairwise comparison data implied by the margin-type condition, is independent of the specific estimators used, and thus applies to various learning algorithms and models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on12
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Preference Ranking Optimization for Human AlignmentFeifan Song, Bowen Yu, Minghao Li, Haiyang Yu et al.AAAI 2024 · 357 citations
- Principled Reinforcement Learning with Human Feedback from Pairwise or K-wise ComparisonsBanghua Zhu, Michael I. Jordan, Jiantao JiaoICML 2023 · 273 citations
- Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHFAnand Siththaranjan, Cassidy Laidlaw, Dylan Hadfield-MenellICLR 2024 · 112 citations
Related papers
- Active Reward Modeling: Adaptive Preference Labeling for Large Language Model AlignmentYunyi Shen, Hao Sun, Jean-Francois TonICML 2025
- The Effect of Modeling Human Rationality Level on Learning Rewards from Multiple Feedback TypesGaurav R. Ghosal, Matthew Zurek, Daniel S. Brown, Anca D. DraganAAAI 2023 · 38 citations
- Neural Contextual Bandits with UCB-based ExplorationDongruo Zhou, Lihong Li, Quanquan GuICML 2020 · 329 citations
- Axioms for AI Alignment from Human FeedbackLuise Ge, Daniel Halpern, Evi Micha, Ariel D. Procaccia et al.NeurIPS 2024 · 64 citations
- A Finite-Time Analysis of Q-Learning with Neural Network Function ApproximationPan Xu, Quanquan GuICML 2020 · 79 citations
