Hindsight Task Relabelling: Experience Replay for Sparse Reward Meta-RL
Charles Packer, Pieter Abbeel, Joseph E. Gonzalez
Abstract
Meta-reinforcement learning (meta-RL) has proven to be a successful framework for leveraging experience from prior tasks to rapidly learn new related tasks, however, current meta-RL approaches struggle to learn in sparse reward environments. Although existing meta-RL algorithms can learn strategies for adapting to new sparse reward tasks, the actual adaptation strategies are learned using hand-shaped reward functions, or require simple environments where random exploration is sufficient to encounter sparse reward. In this paper, we present a formulation of hindsight relabeling for meta-RL, which relabels experience during meta-training to enable learning to learn entirely using sparse reward. We demonstrate the effectiveness of our approach on a suite of challenging sparse reward goal-reaching environments that previously required dense reward during meta-training to solve. Our approach solves these environments using the true sparse reward function, with performance comparable to training with a proxy dense reward function.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0633fd1b-d27b-45e5-8c27-101594cb086aCited by top-tier papers6
- On the Effectiveness of Fine-tuning Versus Meta-reinforcement LearningMandi Zhao, Pieter Abbeel, Stephen JamesNeurIPS 2022 · 43 citations
- Self-Improving Skill Learning for Robust Skill-based Meta-Reinforcement LearningSeungyul Han, Sanghyeon Lee, Sangjun Bae, Yisak ParkICLR 2026 · 5 citations
- Doubly Robust Augmented Transfer for Meta-Reinforcement LearningYuankun Jiang, Nuowen Kan, Chenglin Li, Wenrui Dai et al.NeurIPS 2023 · 3 citations
- Enhanced Meta Reinforcement Learning via Demonstrations in Sparse Reward EnvironmentsDesik Rengarajan, Sapana Chaudhary, Jaewon Kim, Dileep Kalathil et al.NeurIPS 2022 · 2 citations
- Task-Aware Virtual Training: Enhancing Generalization in Meta-Reinforcement Learning for Out-of-Distribution TasksJeongmo Kim, Yisak Park, Minung Kim, Seungyul HanICML 2025
Builds on2
Related papers
- Hindsight Foresight Relabeling for Meta-Reinforcement LearningMichael Wan, Jian Peng, Tanmay GangwaniICLR 2022 · 7 citations
- HMRL: Hyper-Meta Learning for Sparse Reward Reinforcement Learning ProblemYun Hua, Xiangfeng Wang, Bo Jin, Wenhao Li et al.KDD 2021 · 6 citations
- Exploration in Approximate Hyper-State Space for Meta Reinforcement LearningLuisa M. Zintgraf, Leo Feng, Cong Lu, Maximilian Igl et al.ICML 2021 · 45 citations
- Wish you were here: Hindsight Goal Selection for long-horizon dexterous manipulationTodor Davchev, Oleg Olegovich Sushkov, Jean-Baptiste Regli, Stefan Schaal et al.ICLR 2022 · 19 citations
- MetaCURE: Meta Reinforcement Learning with Empowerment-Driven ExplorationJin Zhang, Jianhao Wang, Hao Hu, Tong Chen et al.ICML 2021 · 33 citations
