Model-based Adversarial Meta-Reinforcement Learning
Zichuan Lin, Garrett Thomas, Guangwen Yang, Tengyu Ma
Abstract
Meta-reinforcement learning (meta-RL) aims to learn from multiple training tasks the ability to adapt efficiently to unseen test tasks. Despite the success, existing meta-RL algorithms are known to be sensitive to the task distribution shift. When the test task distribution is different from the training task distribution, the performance may degrade significantly. To address this issue, this paper proposes Model-based Adversarial Meta-Reinforcement Learning (AdMRL), where we aim to minimize the worst-case sub-optimality gap -- the difference between the optimal return and the return that the algorithm achieves after adaptation -- across all tasks in a family of tasks, with a model-based approach. We propose a minimax objective and optimize it by alternating between learning the dynamics model on a fixed task and finding the adversarial task for the current model -- the task for which the policy induced by the model is maximally suboptimal. Assuming the family of tasks is parameterized, we derive a formula for the gradient of the suboptimality with respect to the task parameters via the implicit function theorem, and show how the gradient estimator can be efficiently implemented by the conjugate gradient method and a novel use of the REINFORCE estimator. We evaluate our approach on several continuous control benchmarks and demonstrate its efficacy in the worst-case performance over all tasks, the generalization power to out-of-distribution tasks, and in training and test time sample efficiency, over existing state-of-the-art meta-RL algorithms.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 28377480-0b00-47da-b274-6b2e101ab2e3Cited by top-tier papers16
- Robust Policy Learning over Multiple Uncertainty SetsAnnie Xie, Shagun Sodhani, Chelsea Finn, Joelle Pineau et al.ICML 2022 · 25 citations
- Improving Generalization in Meta-RL with Imaginary Tasks from Latent Dynamics MixtureSuyoung Lee, Sae-Young ChungNeurIPS 2021 · 23 citations
- Distributionally Adaptive Meta Reinforcement LearningAnurag Ajay, Abhishek Gupta, Dibya Ghosh, Sergey Levine et al.NeurIPS 2022 · 21 citations
- Parameterizing Non-Parametric Meta-Reinforcement Learning Tasks via Subtask DecompositionSuyoung Lee, Myungsik Cho, Youngchul SungNeurIPS 2023 · 18 citations
- Max-Min Off-Policy Actor-Critic Method Focusing on Worst-Case Robustness to Model MisspecificationTakumi Tanabe, Rei Sato, Kazuto Fukuchi, Jun Sakuma et al.NeurIPS 2022 · 17 citations
Builds on5
- VariBAD: A Very Good Method for Bayes-Adaptive Deep RL via Meta-LearningLuisa M. Zintgraf, Kyriacos Shiarlis, Maximilian Igl, Sebastian Schulze et al.ICLR 2020 · 315 citations
- Reward-Free Exploration for Reinforcement LearningChi Jin, Akshay Krishnamurthy, Max Simchowitz, Tiancheng YuICML 2020 · 226 citations
- Exploring Model-based Planning with Policy NetworksTingwu Wang, Jimmy BaICLR 2020 · 164 citations
- A Game Theoretic Framework for Model Based Reinforcement LearningAravind Rajeswaran, Igor Mordatch, Vikash KumarICML 2020 · 137 citations
- Improving Generalization in Meta Reinforcement Learning using Learned ObjectivesLouis Kirsch, Sjoerd van Steenkiste, Jürgen SchmidhuberICLR 2020 · 132 citations
Related papers
- Meta-Reinforcement Learning with Universal Policy Adaptation: Provable Near-Optimality under All-task Optimum ComparatorSiyuan Xu, Minghui ZhuNeurIPS 2024 · 8 citations
- Learning Action Translator for Meta Reinforcement Learning on Sparse-Reward TasksYijie Guo, Qiucheng Wu, Honglak LeeAAAI 2022 · 8 citations
- Train Hard, Fight Easy: Robust Meta Reinforcement LearningIdo Greenberg, Shie Mannor, Gal Chechik, Eli A. MeiromNeurIPS 2023 · 15 citations
- Learning to Reweight Imaginary Transitions for Model-Based Reinforcement LearningWenzhen Huang, Qiyue Yin, Junge Zhang, Kaiqi HuangAAAI 2021 · 3 citations
- Offline Meta-Reinforcement Learning with Online Self-SupervisionVitchyr H. Pong, Ashvin Nair, Laura Smith, Catherine Huang et al.ICML 2022 · 78 citations
