PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-training
Kimin Lee, Laura M. Smith, Pieter Abbeel
Abstract
Conveying complex objectives to reinforcement learning (RL) agents can often be difficult, involving meticulous design of reward functions that are sufficiently informative yet easy enough to provide. Human-in-the-loop RL methods allow practitioners to instead interactively teach agents through tailored feedback; however, such approaches have been challenging to scale since human feedback is very expensive. In this work, we aim to make this process more sample- and feedback-efficient. We present an off-policy, interactive RL algorithm that capitalizes on the strengths of both feedback and off-policy learning. Specifically, we learn a reward model by actively querying a teacher's preferences between two clips of behavior and use it to train an agent. To enable off-policy learning, we relabel all the agent's past experience when its reward model changes. We additionally show that pre-training our agents with unsupervised exploration substantially increases the mileage of its queries. We demonstrate that our approach is capable of learning tasks of higher complexity than previously considered by human-in-the-loop methods, including a variety of locomotion and robotic manipulation skills. We also show that our method is able to utilize real-time human feedback to effectively prevent reward exploitation and learn new behaviors that are difficult to specify with standard reward functions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 16c714cd-7e3d-495a-ac66-6a019b7c0339Cited by top-tier papers124
- ImageReward: Learning and Evaluating Human Preferences for Text-to-Image GenerationJiazheng Xu, Xiao Liu, Yuchen Wu, Yuxuan Tong et al.NeurIPS 2023 · 1,310 citations
- Reinforcement Learning for Fine-tuning Text-to-Image Diffusion ModelsYing Fan, Olivia Watkins, Yuqing Du, Hao Liu et al.NeurIPS 2023 · 372 citations
- Preference Ranking Optimization for Human AlignmentFeifan Song, Bowen Yu, Minghao Li, Haiyang Yu et al.AAAI 2024 · 357 citations
- Reward Model Ensembles Help Mitigate OveroptimizationThomas Coste, Usman Anwar, Robert Kirk, David KruegerICLR 2024 · 208 citations
- RoboCLIP: One Demonstration is Enough to Learn Robot PoliciesSumedh Sontakke, Jesse Zhang, Sébastien M. R. Arnold, Karl Pertsch et al.NeurIPS 2023 · 182 citations
Builds on4
- Dynamics-Aware Unsupervised Discovery of SkillsArchit Sharma, Shixiang Gu, Sergey Levine, Vikash Kumar et al.ICLR 2020 · 475 citations
- Behavior From the Void: Unsupervised Active Pre-TrainingHao Liu, Pieter AbbeelNeurIPS 2021 · 258 citations
- State Entropy Maximization with Random Encoders for Efficient ExplorationYounggyo Seo, Lili Chen, Jinwoo Shin, Honglak Lee et al.ICML 2021 · 158 citations
- Avoiding Side Effects in Complex EnvironmentsAlexander Matt Turner, Neale Ratzlaff, Prasad TadepalliNeurIPS 2020 · 40 citations
Related papers
- Provably Feedback-Efficient Reinforcement Learning via Active Reward LearningDingwen Kong, Lin YangNeurIPS 2022 · 19 citations
- Efficient Meta Reinforcement Learning for Preference-based Fast AdaptationZhizhou Ren, Anji Liu, Yitao Liang, Jian Peng et al.NeurIPS 2022 · 11 citations
- Widening the Pipeline in Human-Guided Reinforcement Learning with Explanation and Context-Aware Data AugmentationLin Guan, Mudit Verma, Sihang Guo, Ruohan Zhang et al.NeurIPS 2021 · 57 citations
- Leveraging Sub-Optimal Data for Human-in-the-Loop Reinforcement LearningCalarina Muslimani, Matthew E. TaylorICLR 2025
- Teachable Reinforcement Learning via Advice DistillationOlivia Watkins, Abhishek Gupta, Trevor Darrell, Pieter Abbeel et al.NeurIPS 2021 · 3 citations
