On the Convergence Theory of Debiased Model-Agnostic Meta-Reinforcement Learning
Alireza Fallah, Kristian Georgiev, Aryan Mokhtari, Asuman E. Ozdaglar
摘要
We consider Model-Agnostic Meta-Learning (MAML) methods for Reinforcement Learning (RL) problems, where the goal is to find a policy using data from several tasks represented by Markov Decision Processes (MDPs) that can be updated by one step of stochastic policy gradient for the realized MDP. In particular, using stochastic gradients in MAML update steps is crucial for RL problems since computation of exact gradients requires access to a large number of possible trajectories. For this formulation, we propose a variant of the MAML method, named Stochastic Gradient Meta-Reinforcement Learning (SG-MRL), and study its convergence properties. We derive the iteration and sample complexity of SG-MRL to find an -first-order stationary point, which, to the best of our knowledge, provides the first convergence guarantee for model-agnostic meta-reinforcement learning algorithms. We further show how our results extend to the case where more than one step of stochastic policy gradient method is used at test time. Finally, we empirically compare SG-MRL and MAML in several deep RL environments.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Meta Inverse Constrained Reinforcement Learning: Convergence Guarantee and Generalization AnalysisShicheng Liu, Minghui ZhuICLR 2024 · 被引用 26 次
- A Single-timescale Analysis for Stochastic Approximation with Multiple Coupled SequencesHan Shen, Tianyi ChenNeurIPS 2022 · 被引用 25 次
- Multi-Agent Meta-Reinforcement Learning: Sharper Convergence Rates with Task SimilarityWeichao Mao, Haoran Qiu, Chen Wang, Hubertus Franke 等NeurIPS 2023 · 被引用 17 次
- A Theoretical Understanding of Gradient Bias in Meta-Reinforcement LearningBo Liu, Xidong Feng, Jie Ren, Luo Mai 等NeurIPS 2022 · 被引用 16 次
- Meta-Reinforcement Learning with Universal Policy Adaptation: Provable Near-Optimality under All-task Optimum ComparatorSiyuan Xu, Minghui ZhuNeurIPS 2024 · 被引用 8 次
它引用的顶会 Paper1
相关 Paper
- ES-MAML: Simple Hessian-Free Meta LearningXingyou Song, Wenbo Gao, Yuxiang Yang, Krzysztof Choromanski 等ICLR 2020 · 被引用 128 次
- On the Global Optimality of Model-Agnostic Meta-LearningLingxiao Wang, Qi Cai, Zhuoran Yang, Zhaoran WangICML 2020 · 被引用 48 次
- Task-Robust Model-Agnostic Meta-LearningLiam Collins, Aryan Mokhtari, Sanjay ShakkottaiNeurIPS 2020 · 被引用 66 次
- Model-based Adversarial Meta-Reinforcement LearningZichuan Lin, Garrett Thomas, Guangwen Yang, Tengyu MaNeurIPS 2020 · 被引用 58 次
- Generalization Bounds For Meta-Learning: An Information-Theoretic AnalysisQi Chen, Changjian Shui, Mario MarchandNeurIPS 2021 · 被引用 66 次
