A Theoretical Understanding of Gradient Bias in Meta-Reinforcement Learning
Bo Liu, Xidong Feng, Jie Ren, Luo Mai, Rui Zhu, Haifeng Zhang, Jun Wang, Yaodong Yang
摘要
Gradient-based Meta-RL (GMRL) refers to methods that maintain two-level optimisation procedures wherein the outer-loop meta-learner guides the inner-loop gradient-based reinforcement learner to achieve fast adaptations. In this paper, we develop a unified framework that describes variations of GMRL algorithms and points out that existing stochastic meta-gradient estimators adopted by GMRL are actually biased. Such meta-gradient bias comes from two sources: 1) the compositional bias incurred by the two-level problem structure, which has an upper bound of O 𝐾𝛼 𝐾 σIn |𝜏| -0.5 w.r.t. inner-loop update step 𝐾, learning rate 𝛼, estimate variance σ2 In and sample size |𝜏|, and 2) the multi-step Hessian estimation bias Δ𝐻 due to the use of autodiff, which has a polynomial impact O (𝐾 -1) ( Δ𝐻 ) 𝐾 -1 on the meta-gradient bias. We study tabular MDPs empirically and offer quantitative evidence that testifies our theoretical findings on existing stochastic meta-gradient estimators. Furthermore, we conduct experiments on Iterated Prisoner's Dilemma and Atari games to show how other methods such as off-policy learning and low-bias estimator can help fix the gradient bias for GMRL algorithms in general. * Equal contribution, the order is determined by flipping a coin. See Appendix J for more details. † Corresponding author. 36th Conference on Neural Information Processing Systems (NeurIPS 2022).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- EvIL: Evolution Strategies for Generalisable Imitation LearningSilvia Sapora, Gokul Swamy, Chris Lu, Yee Whye Teh 等ICML 2024 · 被引用 10 次
- MetaCARD: Meta-Reinforcement Learning with Task Uncertainty Feedback via Decoupled Context-Aware Reward and Dynamics ComponentsMin Wang, Xin Li, Leiji Zhang, Mingzhong WangAAAI 2024 · 被引用 6 次
- GEAR: A GPU-Centric Experience Replay System for Large Reinforcement Learning ModelsHanjing Wang, Man-Kit Sit, Congjie He, Ying Wen 等ICML 2023 · 被引用 5 次
- Meta-Learning Objectives for Preference OptimizationCarlo Alfano, Silvia Sapora, Jakob N. Foerster, Patrick Rebeschini 等NeurIPS 2025 · 被引用 3 次
- CERTAIN: Context Uncertainty-aware One-Shot Adaptation for Context-based Offline Meta Reinforcement LearningHongtu Zhou, Ruiling Yang, Yakun Zhu, Haoqi Zhao 等ICML 2025
它引用的顶会 Paper13
- Learning to Utilize Shaping Rewards: A New Approach of Reward ShapingYujing Hu, Weixun Wang, Hangtian Jia, Yixiang Wang 等NeurIPS 2020 · 被引用 256 次
- Discovering Reinforcement Learning AlgorithmsJunhyuk Oh, Matteo Hessel, Wojciech M. Czarnecki, Zhongwen Xu 等NeurIPS 2020 · 被引用 154 次
- Improving Generalization in Meta Reinforcement Learning using Learned ObjectivesLouis Kirsch, Sjoerd van Steenkiste, Jürgen SchmidhuberICLR 2020 · 被引用 132 次
- A Self-Tuning Actor-Critic AlgorithmTom Zahavy, Zhongwen Xu, Vivek Veeriah, Matteo Hessel 等NeurIPS 2020 · 被引用 106 次
- Meta-Gradient Reinforcement Learning with an Objective Discovered OnlineZhongwen Xu, Hado Philip van Hasselt, Matteo Hessel, Junhyuk Oh 等NeurIPS 2020 · 被引用 90 次
相关 Paper
- Unifying Gradient Estimators for Meta-Reinforcement Learning via Off-Policy EvaluationYunhao Tang, Tadashi Kozuno, Mark Rowland, Rémi Munos 等NeurIPS 2021 · 被引用 9 次
- Biased Gradient Estimate with Drastic Variance Reduction for Meta Reinforcement LearningYunhao TangICML 2022 · 被引用 7 次
- Distributional Meta-Gradient Reinforcement LearningHaiyan Yin, Shuicheng Yan, Zhongwen XuICLR 2023
- On the Convergence Theory of Debiased Model-Agnostic Meta-Reinforcement LearningAlireza Fallah, Kristian Georgiev, Aryan Mokhtari, Asuman E. OzdaglarNeurIPS 2021 · 被引用 31 次
- Statistically Efficient Off-Policy Policy GradientsNathan Kallus, Masatoshi UeharaICML 2020 · 被引用 43 次
