Enhanced Meta Reinforcement Learning via Demonstrations in Sparse Reward Environments
Desik Rengarajan, Sapana Chaudhary, Jaewon Kim, Dileep Kalathil, Srinivas Shakkottai
Abstract
Meta reinforcement learning (Meta-RL) is an approach wherein the experience gained from solving a variety of tasks is distilled into a meta-policy. The metapolicy, when adapted over only a small (or just a single) number of steps, is able to perform near-optimally on a new, related task. However, a major challenge to adopting this approach to solve real-world problems is that they are often associated with sparse reward functions that only indicate whether a task is completed partially or fully. We consider the situation where some data, possibly generated by a suboptimal agent, is available for each task. We then develop a class of algorithms entitled Enhanced Meta-RL using Demonstrations (EMRLD) that exploit this information-even if sub-optimal-to obtain guidance during training. We show how EMRLD jointly utilizes RL and supervised learning over the offline data to generate a meta-policy that demonstrates monotone performance improvements. We also develop a warm started variant called EMRLD-WS that is particularly efficient for sub-optimal demonstration data. Finally, we show that our EMRLD algorithms significantly outperform existing approaches in a variety of sparse reward environments, including that of a mobile robot.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4d8fbe74-3607-414b-8c0a-017c44ce7d30Cited by top-tier papers1
Ask how each one uses itBuilds on3
- Reinforcement Learning with Sparse Rewards using Guidance from Offline DemonstrationDesik Rengarajan, Gargi Vaidya, Akshay Sarvesh, Dileep M. Kalathil et al.ICLR 2022 · 86 citations
- Watch, Try, Learn: Meta-Learning from Demonstrations and RewardsAllan Zhou, Eric Jang, Daniel Kappler, Alexander Herzog et al.ICLR 2020 · 53 citations
- Hindsight Task Relabelling: Experience Replay for Sparse Reward Meta-RLCharles Packer, Pieter Abbeel, Joseph E. GonzalezNeurIPS 2021 · 22 citations
Related papers
- Offline Meta-Reinforcement Learning with Online Self-SupervisionVitchyr H. Pong, Ashvin Nair, Laura Smith, Catherine Huang et al.ICML 2022 · 78 citations
- HMRL: Hyper-Meta Learning for Sparse Reward Reinforcement Learning ProblemYun Hua, Xiangfeng Wang, Bo Jin, Wenhao Li et al.KDD 2021 · 6 citations
- Skill-based Meta-Reinforcement LearningTaewook Nam, Shao-Hua Sun, Karl Pertsch, Sung Ju Hwang et al.ICLR 2022 · 55 citations
- Learning Action Translator for Meta Reinforcement Learning on Sparse-Reward TasksYijie Guo, Qiucheng Wu, Honglak LeeAAAI 2022 · 8 citations
- Jump-Start Reinforcement LearningIkechukwu Uchendu, Ted Xiao, Yao Lu, Banghua Zhu et al.ICML 2023 · 158 citations
