Reset-Free Lifelong Learning with Skill-Space Planning
Kevin Lu, Aditya Grover, Pieter Abbeel, Igor Mordatch
Abstract
The objective of lifelong reinforcement learning (RL) is to optimize agents which can continuously adapt and interact in changing environments. However, current RL approaches fail drastically when environments are non-stationary and interactions are non-episodic. We propose Lifelong Skill Planning (LiSP), an algorithmic framework for non-episodic lifelong RL based on planning in an abstract space of higher-order skills. We learn the skills in an unsupervised manner using intrinsic rewards and plan over the learned skills using a learned dynamics model. Moreover, our framework permits skill discovery even from offline data, thereby reducing the need for excessive real-world interactions. We demonstrate empirically that LiSP successfully enables long-horizon planning and learns agents that can avoid catastrophic failures even in challenging non-stationary and non-episodic environments derived from gridworld and MuJoCo benchmarks 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e3ce1138-9bf4-47c3-ae3e-2da3dc9c88b5Cited by top-tier papers15
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
- Online Decision TransformerQinqing Zheng, Amy Zhang, Aditya GroverICML 2022 · 256 citations
- Autonomous Reinforcement Learning via Subgoal CurriculaArchit Sharma, Abhishek Gupta, Sergey Levine, Karol Hausman et al.NeurIPS 2021 · 41 citations
- Beyond Reward: Offline Preference-guided Policy OptimizationYachen Kang, Diyuan Shi, Jinxin Liu, Li He et al.ICML 2023 · 41 citations
- Autonomous Reinforcement Learning: Formalism and BenchmarkingArchit Sharma, Kelvin Xu, Nikhil Sardana, Abhishek Gupta et al.ICLR 2022 · 39 citations
Builds on12
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon et al.NeurIPS 2020 · 989 citations
- An Optimistic Perspective on Offline Reinforcement LearningRishabh Agarwal, Dale Schuurmans, Mohammad NorouziICML 2020 · 568 citations
- Dynamics-Aware Unsupervised Discovery of SkillsArchit Sharma, Shixiang Gu, Sergey Levine, Vikash Kumar et al.ICLR 2020 · 475 citations
- The Ingredients of Real World Robotic Reinforcement LearningHenry Zhu, Justin Yu, Abhishek Gupta, Dhruv Shah et al.ICLR 2020 · 202 citations
Related papers
- Meta-learning Parameterized SkillsHaotian Fu, Shangqun Yu, Saket Tiwari, Michael Littman et al.ICML 2023 · 8 citations
- One After Another: Learning Incremental Skills for a Changing WorldNur Muhammad (Mahi) Shafiullah, Lerrel PintoICLR 2022 · 15 citations
- Deep Reinforcement Learning amidst Continual Structured Non-StationarityAnnie Xie, James Harrison, Chelsea FinnICML 2021 · 43 citations
- Foundation Policies with Hilbert RepresentationsSeohong Park, Tobias Kreiman, Sergey LevineICML 2024 · 72 citations
- Learning Temporally AbstractWorld Models without Online ExperimentationBenjamin Freed, Siddarth Venkatraman, Guillaume Adrien Sartoretti, Jeff Schneider et al.ICML 2023 · 7 citations
