MetaCURE: Meta Reinforcement Learning with Empowerment-Driven Exploration
Jin Zhang, Jianhao Wang, Hao Hu, Tong Chen, Yingfeng Chen, Changjie Fan, Chongjie Zhang
摘要
Meta reinforcement learning (meta-RL) extracts knowledge from previous tasks and achieves fast adaptation to new tasks. Despite recent progress, efficient exploration in meta-RL remains a key challenge in sparse-reward tasks, as it requires quickly finding informative task-relevant experiences in both meta-training and adaptation. To address this challenge, we explicitly model an exploration policy learning problem for meta-RL, which is separated from exploitation policy learning, and introduce a novel empowerment-driven exploration objective, which aims to maximize information gain for task identification. We derive a corresponding intrinsic reward and develop a new off-policy meta-RL framework, which efficiently learns separate context-aware exploration and exploitation policies by sharing the knowledge of task inference. Experimental evaluation shows that our meta-RL method significantly outperforms state-of-the-art baselines on various sparse-reward MuJoCo locomotion tasks and more complex sparse-reward Meta-World tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- Bridging State and History Representations: Understanding Self-Predictive RLTianwei Ni, Benjamin Eysenbach, Erfan Seyedsalehi, Michel Ma 等ICLR 2024 · 被引用 50 次
- On the Effectiveness of Fine-tuning Versus Meta-reinforcement LearningMandi Zhao, Pieter Abbeel, Stephen JamesNeurIPS 2022 · 被引用 43 次
- Procedural generalization by planning with self-supervised world modelsAnkesh Anand, Jacob C. Walker, Yazhe Li, Eszter Vértes 等ICLR 2022 · 被引用 34 次
- Context Shift Reduction for Offline Meta-Reinforcement LearningYunkai Gao, Rui Zhang, Jiaming Guo, Fan Wu 等NeurIPS 2023 · 被引用 30 次
- Robust Policy Learning over Multiple Uncertainty SetsAnnie Xie, Shagun Sodhani, Chelsea Finn, Joelle Pineau 等ICML 2022 · 被引用 25 次
它引用的顶会 Paper3
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 被引用 1,852 次
- VariBAD: A Very Good Method for Bayes-Adaptive Deep RL via Meta-LearningLuisa M. Zintgraf, Kyriacos Shiarlis, Maximilian Igl, Sebastian Schulze 等ICLR 2020 · 被引用 315 次
- Context-aware Dynamics Model for Generalization in Model-Based Reinforcement LearningKimin Lee, Younggyo Seo, Seunghyun Lee, Honglak Lee 等ICML 2020 · 被引用 158 次
相关 Paper
- HMRL: Hyper-Meta Learning for Sparse Reward Reinforcement Learning ProblemYun Hua, Xiangfeng Wang, Bo Jin, Wenhao Li 等KDD 2021 · 被引用 6 次
- Doubly Robust Augmented Transfer for Meta-Reinforcement LearningYuankun Jiang, Nuowen Kan, Chenglin Li, Wenrui Dai 等NeurIPS 2023 · 被引用 3 次
- Learning Action Translator for Meta Reinforcement Learning on Sparse-Reward TasksYijie Guo, Qiucheng Wu, Honglak LeeAAAI 2022 · 被引用 8 次
- MetaCARD: Meta-Reinforcement Learning with Task Uncertainty Feedback via Decoupled Context-Aware Reward and Dynamics ComponentsMin Wang, Xin Li, Leiji Zhang, Mingzhong WangAAAI 2024 · 被引用 6 次
- Decoupling Exploration and Exploitation for Meta-Reinforcement Learning without SacrificesEvan Zheran Liu, Aditi Raghunathan, Percy Liang, Chelsea FinnICML 2021 · 被引用 80 次
