Imitating Past Successes can be Very Suboptimal
Benjamin Eysenbach, Soumith Udatha, Russ Salakhutdinov, Sergey Levine
Abstract
Prior work has proposed a simple strategy for reinforcement learning (RL): label experience with the outcomes achieved in that experience, and then imitate the relabeled experience. These outcome-conditioned imitation learning methods are appealing because of their simplicity, strong performance, and close ties with supervised learning. However, it remains unclear how these methods relate to the standard RL objective, reward maximization. In this paper, we formally relate outcome-conditioned imitation learning to reward maximization, drawing a precise relationship between the learned policy and Q-values and explaining the close connections between these methods and prior EM-based policy search methods. This analysis shows that existing outcome-conditioned imitation learning methods do not necessarily improve the policy, but a simple modification results in a method that does guarantee policy improvement, under some assumptions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 80047bec-bbbc-4e15-8919-1874924cfcc7Cited by top-tier papers11
- Contrastive Learning as Goal-Conditioned Reinforcement LearningBenjamin Eysenbach, Tianjun Zhang, Sergey Levine, Ruslan SalakhutdinovNeurIPS 2022 · 331 citations
- A Policy-Guided Imitation Approach for Offline Reinforcement LearningHaoran Xu, Li Jiang, Jianxiong Li, Xianyuan ZhanNeurIPS 2022 · 86 citations
- Offline Multi-Agent Reinforcement Learning with Implicit Global-to-Local Value RegularizationXiangsen Wang, Haoran Xu, Yinan Zheng, Xianyuan ZhanNeurIPS 2023 · 65 citations
- Closing the Gap between TD Learning and Supervised Learning - A Generalisation Point of ViewRaj Ghugare, Matthieu Geist, Glen Berseth, Benjamin EysenbachICLR 2024 · 28 citations
- Score Models for Offline Goal-Conditioned Reinforcement LearningHarshit Sikchi, Rohan Chitnis, Ahmed Touati, Alborz Geramifard et al.ICLR 2024 · 16 citations
Builds on7
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 1,402 citations
- Learning to Reach Goals via Iterated Supervised LearningDibya Ghosh, Abhishek Gupta, Ashwin Reddy, Justin Fu et al.ICLR 2021 · 222 citations
- Rewriting History with Inverse RL: Hindsight Inference for Policy ImprovementBen Eysenbach, Xinyang Geng, Sergey Levine, Ruslan SalakhutdinovNeurIPS 2020 · 96 citations
- Generalized Hindsight for Reinforcement LearningAlexander C. Li, Lerrel Pinto, Pieter AbbeelNeurIPS 2020 · 81 citations
Related papers
- SEABO: A Simple Search-Based Method for Offline Imitation LearningJiafei Lyu, Xiaoteng Ma, Le Wan, Runze Liu et al.ICLR 2024 · 17 citations
- Imitation Learning by Reinforcement LearningKamil CiosekICLR 2022 · 22 citations
- Representation Matters: Offline Pretraining for Sequential Decision MakingMengjiao Yang, Ofir NachumICML 2021 · 126 citations
- Hindsight Task Relabelling: Experience Replay for Sparse Reward Meta-RLCharles Packer, Pieter Abbeel, Joseph E. GonzalezNeurIPS 2021 · 22 citations
- Adversarial Intrinsic Motivation for Reinforcement LearningIshan Durugkar, Mauricio Tec, Scott Niekum, Peter StoneNeurIPS 2021 · 61 citations
