Exploration in Approximate Hyper-State Space for Meta Reinforcement Learning
Luisa M. Zintgraf, Leo Feng, Cong Lu, Maximilian Igl, Kristian Hartikainen, Katja Hofmann, Shimon Whiteson
摘要
Meta-learning is a powerful tool for learning policies that can adapt efficiently when deployed in new tasks. If however the meta-training tasks have sparse rewards, the need for exploration during meta-training is exacerbated given that the agent has to explore and learn across many tasks. We show that current meta-learning methods can fail catastrophically in such environments. To address this problem, we propose HyperX, a novel method for meta-learning in sparse reward tasks. Using novel reward bonuses for meta-training, we incentivise the agent to explore in approximate hyper-state space, i.e., the joint state and approximate belief space, where the beliefs are over tasks. We show empirically that these bonuses allow an agent to successfully learn to solve sparse reward tasks where existing meta-learning methods fail.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- Augmented World Models Facilitate Zero-Shot Dynamics Generalization From a Single Offline EnvironmentPhilip J. Ball, Cong Lu, Jack Parker-Holder, Stephen J. RobertsICML 2021 · 被引用 55 次
- Bridging State and History Representations: Understanding Self-Predictive RLTianwei Ni, Benjamin Eysenbach, Erfan Seyedsalehi, Michel Ma 等ICLR 2024 · 被引用 50 次
- On the Effectiveness of Fine-tuning Versus Meta-reinforcement LearningMandi Zhao, Pieter Abbeel, Stephen JamesNeurIPS 2022 · 被引用 43 次
- Dynamics Generalisation in Reinforcement Learning via Adaptive Context-Aware PoliciesMichael Beukman, Devon Jarvis, Richard Klein, Steven James 等NeurIPS 2023 · 被引用 29 次
- Meta-Reinforcement Learning Based on Self-Supervised Task Representation LearningMingyang Wang, Zhenshan Bing, Xiangtong Yao, Shuai Wang 等AAAI 2023 · 被引用 22 次
它引用的顶会 Paper3
- VariBAD: A Very Good Method for Bayes-Adaptive Deep RL via Meta-LearningLuisa M. Zintgraf, Kyriacos Shiarlis, Maximilian Igl, Sebastian Schulze 等ICLR 2020 · 被引用 315 次
- Conservative Uncertainty Estimation By Fitting Prior NetworksKamil Ciosek, Vincent Fortuin, Ryota Tomioka, Katja Hofmann 等ICLR 2020 · 被引用 65 次
- Meta-trained agents implement Bayes-optimal agentsVladimir Mikulik, Grégoire Delétang, Tom McGrath, Tim Genewein 等NeurIPS 2020 · 被引用 56 次
相关 Paper
- HMRL: Hyper-Meta Learning for Sparse Reward Reinforcement Learning ProblemYun Hua, Xiangfeng Wang, Bo Jin, Wenhao Li 等KDD 2021 · 被引用 6 次
- Hindsight Task Relabelling: Experience Replay for Sparse Reward Meta-RLCharles Packer, Pieter Abbeel, Joseph E. GonzalezNeurIPS 2021 · 被引用 22 次
- MetaCURE: Meta Reinforcement Learning with Empowerment-Driven ExplorationJin Zhang, Jianhao Wang, Hao Hu, Tong Chen 等ICML 2021 · 被引用 33 次
- Watch, Try, Learn: Meta-Learning from Demonstrations and RewardsAllan Zhou, Eric Jang, Daniel Kappler, Alexander Herzog 等ICLR 2020 · 被引用 53 次
- Enhanced Meta Reinforcement Learning via Demonstrations in Sparse Reward EnvironmentsDesik Rengarajan, Sapana Chaudhary, Jaewon Kim, Dileep Kalathil 等NeurIPS 2022 · 被引用 2 次
