Learning Action Translator for Meta Reinforcement Learning on Sparse-Reward Tasks
Yijie Guo, Qiucheng Wu, Honglak Lee
摘要
Meta reinforcement learning (meta-RL) aims to learn a policy solving a set of training tasks simultaneously and quickly adapting to new tasks. It requires massive amounts of data drawn from training tasks to infer the common structure shared among tasks. Without heavy reward engineering, the sparse rewards in long-horizon tasks exacerbate the problem of sample efficiency in meta-RL. Another challenge in meta-RL is the discrepancy of difficulty level among tasks, which might cause one easy task dominating learning of the shared policy and thus preclude policy adaptation to new tasks. This work introduces a novel objective function to learn an action translator among training tasks. We theoretically verify that the value of the transferred policy with the action translator can be close to the value of the source policy and our objective function (approximately) upper bounds the value difference. We propose to combine the action translator with context-based meta-RL algorithms for better data collection and moreefficient exploration during meta-training. Our approach em-pirically improves the sample efficiency and performance ofmeta-RL algorithms on sparse-reward tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Meta-Reinforcement Learning Based on Self-Supervised Task Representation LearningMingyang Wang, Zhenshan Bing, Xiangtong Yao, Shuai Wang 等AAAI 2023 · 被引用 22 次
- Analyzing Generalization in Policy Networks: A Case Study with the Double-Integrator SystemRuining Zhang, Haoran Han, Maolong Lv, Qisong Yang 等AAAI 2024 · 被引用 5 次
- Learning Task Belief Similarity with Latent Dynamics for Meta-Reinforcement LearningMenglong Zhang, Fuyuan Qian, Quanying LiuICLR 2025
- SRSA: Skill Retrieval and Adaptation for Robotic Assembly TasksYijie Guo, Bingjie Tang, Iretiayo Akinola, Dieter Fox 等ICLR 2025
它引用的顶会 Paper3
- Invariant Causal Prediction for Block MDPsAmy Zhang, Clare Lyle, Shagun Sodhani, Angelos Filos 等ICML 2020 · 被引用 153 次
- Self-Imitation Learning via Generalized Lower Bound Q-learningYunhao TangNeurIPS 2020 · 被引用 30 次
- Learning Robust State Abstractions for Hidden-Parameter Block MDPsAmy Zhang, Shagun Sodhani, Khimya Khetarpal, Joelle PineauICLR 2021 · 被引用 5 次
相关 Paper
- HMRL: Hyper-Meta Learning for Sparse Reward Reinforcement Learning ProblemYun Hua, Xiangfeng Wang, Bo Jin, Wenhao Li 等KDD 2021 · 被引用 6 次
- MetaCURE: Meta Reinforcement Learning with Empowerment-Driven ExplorationJin Zhang, Jianhao Wang, Hao Hu, Tong Chen 等ICML 2021 · 被引用 33 次
- Towards Effective Context for Meta-Reinforcement Learning: an Approach based on Contrastive LearningHaotian Fu, Hongyao Tang, Jianye Hao, Chen Chen 等AAAI 2021 · 被引用 61 次
- Doubly Robust Augmented Transfer for Meta-Reinforcement LearningYuankun Jiang, Nuowen Kan, Chenglin Li, Wenrui Dai 等NeurIPS 2023 · 被引用 3 次
- Enhanced Meta Reinforcement Learning via Demonstrations in Sparse Reward EnvironmentsDesik Rengarajan, Sapana Chaudhary, Jaewon Kim, Dileep Kalathil 等NeurIPS 2022 · 被引用 2 次
