Principled Reinforcement Learning with Human Feedback from Pairwise or K-wise Comparisons
Banghua Zhu, Michael I. Jordan, Jiantao Jiao
摘要
We provide a theoretical framework for Reinforcement Learning with Human Feedback (RLHF). Our analysis shows that when the true reward function is linear, the widely used maximum likelihood estimator (MLE) converges under both the Bradley-Terry-Luce (BTL) model and the Plackett-Luce (PL) model. However, we show that when training a policy based on the learned reward model, MLE fails while a pessimistic MLE provides policies with improved performance under certain coverage assumptions. Additionally, we demonstrate that under the PL model, the true MLE and an alternative MLE that splits the -wise comparison into pairwise comparisons both converge. Moreover, the true MLE is asymptotically more efficient. Our results validate the empirical success of existing RLHF algorithms in InstructGPT and provide new insights for algorithm design. Furthermore, our results unify the problem of RLHF and max-entropy Inverse Reinforcement Learning (IRL), and provide the first sample complexity bound for max-entropy IRL.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper131
- Preference Ranking Optimization for Human AlignmentFeifan Song, Bowen Yu, Minghao Li, Haiyang Yu 等AAAI 2024 · 被引用 357 次
- Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-constraintWei Xiong, Hanze Dong, Chenlu Ye, Ziqi Wang 等ICML 2024 · 被引用 346 次
- MEM1: Learning to Synergize Memory and Reasoning for Efficient Long-Horizon AgentsZijian Zhou, Ao Qu, Zhaoxuan Wu, Sunghwan Kim 等ICLR 2026 · 被引用 223 次
- ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language ModelsZiniu Li, Tian Xu, Yushun Zhang, Zhihang Lin 等ICML 2024 · 被引用 165 次
- Text2Reward: Reward Shaping with Language Models for Reinforcement LearningTianbao Xie, Siheng Zhao, Chen Henry Wu, Yitao Liu 等ICLR 2024 · 被引用 142 次
它引用的顶会 Paper16
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 被引用 1,402 次
- Scaling Laws for Reward Model OveroptimizationLeo Gao, John Schulman, Jacob HiltonICML 2023 · 被引用 963 次
- Is Pessimism Provably Efficient for Offline RL?Ying Jin, Zhuoran Yang, Zhaoran WangICML 2021 · 被引用 419 次
相关 Paper
- Axioms for AI Alignment from Human FeedbackLuise Ge, Daniel Halpern, Evi Micha, Ariel D. Procaccia 等NeurIPS 2024 · 被引用 64 次
- Greedy Sampling Is Provably Efficient For RLHFDi Wu, Chengshuai Shi, Jing Yang, Cong ShenNeurIPS 2025 · 被引用 11 次
- Improving LLM General Preference Alignment via Optimistic Online Mirror DescentYuheng Zhang, Dian Yu, Tao Ge, Linfeng Song 等NeurIPS 2025 · 被引用 27 次
- Online Iterative Reinforcement Learning from Human Feedback with General Preference ModelChenlu Ye, Wei Xiong, Yuheng Zhang, Hanze Dong 等NeurIPS 2024 · 被引用 60 次
- Iterative Nash Policy Optimization: Aligning LLMs with General Preferences via No-Regret LearningYuheng Zhang, Dian Yu, Baolin Peng, Linfeng Song 等ICLR 2025
