Inverse Preference Learning: Preference-based RL without a Reward Function
Joey Hejna, Dorsa Sadigh
摘要
Reward functions are difficult to design and often hard to align with human intent. Preference-based Reinforcement Learning (RL) algorithms address these problems by learning reward functions from human feedback. However, the majority of preference-based RL methods naïvely combine supervised reward models with off-the-shelf RL algorithms. Contemporary approaches have sought to improve performance and query complexity by using larger and more complex reward architectures such as transformers. Instead of using highly complex architectures, we develop a new and parameter-efficient algorithm, Inverse Preference Learning (IPL), specifically designed for learning from offline preference data. Our key insight is that for a fixed policy, the -function encodes all information about the reward function, effectively making them interchangeable. Using this insight, we completely eliminate the need for a learned reward function. Our resulting algorithm is simpler and more parameter-efficient. Across a suite of continuous control and robotics benchmarks, IPL attains competitive performance compared to more complex approaches that leverage transformer-based and non-Markovian reward functions while having fewer algorithmic hyperparameters and learned network parameters. Our code is publicly released.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper43
- A Dense Reward View on Aligning Text-to-Image Diffusion with PreferenceShentao Yang, Tianqi Chen, Mingyuan ZhouICML 2024 · 被引用 53 次
- Contrastive Preference Learning: Learning from Human Feedback without Reinforcement LearningJoey Hejna, Rafael Rafailov, Harshit Sikchi, Chelsea Finn 等ICLR 2024 · 被引用 37 次
- Imitating Language via Scalable Inverse Reinforcement LearningMarkus Wulfmeier, Michael Bloesch, Nino Vieillard, Arun Ahuja 等NeurIPS 2024 · 被引用 26 次
- Flow to Better: Offline Preference-based Reinforcement Learning via Preferred Trajectory GenerationZhilong Zhang, Yihao Sun, Junyin Ye, Tian-Shuo Liu 等ICLR 2024 · 被引用 23 次
- Optimal Design for Human Preference ElicitationSubhojyoti Mukherjee, Anusha Lalitha, Kousha Kalantari, Aniket Deshmukh 等NeurIPS 2024 · 被引用 20 次
它引用的顶会 Paper16
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 被引用 1,402 次
- PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-trainingKimin Lee, Laura M. Smith, Pieter AbbeelICML 2021 · 被引用 380 次
- Implementation Matters in Deep RL: A Case Study on PPO and TRPOLogan Engstrom, Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras 等ICLR 2020 · 被引用 305 次
相关 Paper
- Preference Transformer: Modeling Human Preferences using Transformers for RLChangyeon Kim, Jongjin Park, Jinwoo Shin, Honglak Lee 等ICLR 2023 · 被引用 4 次
- Policy-labeled Preference Learning: Is Preference Enough for RLHF?Taehyun Cho, Seokhun Ju, Seungyub Han, Dohyeong Kim 等ICML 2025
- Inverse Reinforcement Learning in a Continuous State Space with Formal GuaranteesGregory Dexter, Kevin Bello, Jean HonorioNeurIPS 2021 · 被引用 9 次
- A Unified Linear Programming Framework for Offline Reward Learning from Human Demonstrations and FeedbackKihyun Kim, Jiawei Zhang, Asuman E. Ozdaglar, Pablo A. ParriloICML 2024 · 被引用 2 次
- Is Inverse Reinforcement Learning Harder than Standard Reinforcement Learning? A Theoretical PerspectiveLei Zhao, Mengdi Wang, Yu BaiICML 2024 · 被引用 3 次
