Policy-labeled Preference Learning: Is Preference Enough for RLHF?
Taehyun Cho, Seokhun Ju, Seungyub Han, Dohyeong Kim, Kyungjae Lee, Jungwoo Lee
摘要
To design rewards that align with human goals, Reinforcement Learning from Human Feedback (RLHF) has emerged as a prominent technique for learning reward functions from human preferences and optimizing policies via reinforcement learning algorithms. However, existing RLHF methods often misinterpret trajectories as being generated by an optimal policy, causing inaccurate likelihood estimation and suboptimal learning. Inspired by Direct Preference Optimization framework which directly learns optimal policy without explicit reward, we propose policy-labeled preference learning (PPL), to resolve likelihood mismatch issues by modeling human preferences with regret, which reflects behavior policy information. We also provide a contrastive KL regularization, derived from regret-based principles, to enhance RLHF in sequential decision making. Experiments in high-dimensional continuous control tasks demonstrate PPL's significant improvements in offline RLHF performance and its effectiveness in online settings. For more information, visit our project page: https://jjush.github. io/PPL/ .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Pareto Optimal Risk-Agnostic Distributional Bandits with Heavy-Tail RewardsKyungjae Lee, Dohyeong Kim, Taehyun Cho, Chaeyeon Kim 等NeurIPS 2025
- PAWS: Preference Learning with Advantage-Weighted SegmentsAleksandar Taranovic, Onur Celik, Niklas Freymuth, Ge Li 等ICML 2026
- A Regret Minimization Framework on Preference Learning in Large Language ModelsSuhwan Kim, Taehyun Cho, Youngsoo Jang, Geon-Hyeong Kim 等ICML 2026
它引用的顶会 Paper21
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 被引用 1,402 次
- PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-trainingKimin Lee, Laura M. Smith, Pieter AbbeelICML 2021 · 被引用 380 次
- Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-constraintWei Xiong, Hanze Dong, Chenlu Ye, Ziqi Wang 等ICML 2024 · 被引用 346 次
- Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence ConstraintsChaoqi Wang, Yibo Jiang, Chenghao Yang, Han Liu 等ICLR 2024 · 被引用 173 次
相关 Paper
- Contrastive Preference Learning: Learning from Human Feedback without Reinforcement LearningJoey Hejna, Rafael Rafailov, Harshit Sikchi, Chelsea Finn 等ICLR 2024 · 被引用 37 次
- Online Iterative Reinforcement Learning from Human Feedback with General Preference ModelChenlu Ye, Wei Xiong, Yuheng Zhang, Hanze Dong 等NeurIPS 2024 · 被引用 60 次
- Direct Preference-based Policy Optimization without Reward ModelingGaon An, Junhyeok Lee, Xingdong Zuo, Norio Kosaka 等NeurIPS 2023 · 被引用 61 次
- PILAF: Optimal Human Preference Sampling for Reward ModelingYunzhen Feng, Ariel Kwiatkowski, Kunhao Zheng, Julia Kempe 等ICML 2025
- Reward Model Learning vs. Direct Policy Optimization: A Comparative Analysis of Learning from Human PreferencesAndi Nika, Debmalya Mandal, Parameswaran Kamalaruban, Georgios Tzannetos 等ICML 2024 · 被引用 22 次
