A Theoretical Understanding of Gradient Bias in Meta-Reinforcement Learning
Bo Liu, Xidong Feng, Jie Ren, Luo Mai, Rui Zhu, Haifeng Zhang, Jun Wang, Yaodong Yang
Abstract
Gradient-based Meta-RL (GMRL) refers to methods that maintain two-level optimisation procedures wherein the outer-loop meta-learner guides the inner-loop gradient-based reinforcement learner to achieve fast adaptations. In this paper, we develop a unified framework that describes variations of GMRL algorithms and points out that existing stochastic meta-gradient estimators adopted by GMRL are actually biased. Such meta-gradient bias comes from two sources: 1) the compositional bias incurred by the two-level problem structure, which has an upper bound of O 𝐾𝛼 𝐾 σIn |𝜏| -0.5 w.r.t. inner-loop update step 𝐾, learning rate 𝛼, estimate variance σ2 In and sample size |𝜏|, and 2) the multi-step Hessian estimation bias Δ𝐻 due to the use of autodiff, which has a polynomial impact O (𝐾 -1) ( Δ𝐻 ) 𝐾 -1 on the meta-gradient bias. We study tabular MDPs empirically and offer quantitative evidence that testifies our theoretical findings on existing stochastic meta-gradient estimators. Furthermore, we conduct experiments on Iterated Prisoner's Dilemma and Atari games to show how other methods such as off-policy learning and low-bias estimator can help fix the gradient bias for GMRL algorithms in general. * Equal contribution, the order is determined by flipping a coin. See Appendix J for more details. † Corresponding author. 36th Conference on Neural Information Processing Systems (NeurIPS 2022).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- EvIL: Evolution Strategies for Generalisable Imitation LearningSilvia Sapora, Gokul Swamy, Chris Lu, Yee Whye Teh et al.ICML 2024 · 10 citations
- MetaCARD: Meta-Reinforcement Learning with Task Uncertainty Feedback via Decoupled Context-Aware Reward and Dynamics ComponentsMin Wang, Xin Li, Leiji Zhang, Mingzhong WangAAAI 2024 · 6 citations
- GEAR: A GPU-Centric Experience Replay System for Large Reinforcement Learning ModelsHanjing Wang, Man-Kit Sit, Congjie He, Ying Wen et al.ICML 2023 · 5 citations
- Meta-Learning Objectives for Preference OptimizationCarlo Alfano, Silvia Sapora, Jakob N. Foerster, Patrick Rebeschini et al.NeurIPS 2025 · 3 citations
- CERTAIN: Context Uncertainty-aware One-Shot Adaptation for Context-based Offline Meta Reinforcement LearningHongtu Zhou, Ruiling Yang, Yakun Zhu, Haoqi Zhao et al.ICML 2025
Builds on13
- Learning to Utilize Shaping Rewards: A New Approach of Reward ShapingYujing Hu, Weixun Wang, Hangtian Jia, Yixiang Wang et al.NeurIPS 2020 · 256 citations
- Discovering Reinforcement Learning AlgorithmsJunhyuk Oh, Matteo Hessel, Wojciech M. Czarnecki, Zhongwen Xu et al.NeurIPS 2020 · 154 citations
- Improving Generalization in Meta Reinforcement Learning using Learned ObjectivesLouis Kirsch, Sjoerd van Steenkiste, Jürgen SchmidhuberICLR 2020 · 132 citations
- A Self-Tuning Actor-Critic AlgorithmTom Zahavy, Zhongwen Xu, Vivek Veeriah, Matteo Hessel et al.NeurIPS 2020 · 106 citations
- Meta-Gradient Reinforcement Learning with an Objective Discovered OnlineZhongwen Xu, Hado Philip van Hasselt, Matteo Hessel, Junhyuk Oh et al.NeurIPS 2020 · 90 citations
Related papers
- Unifying Gradient Estimators for Meta-Reinforcement Learning via Off-Policy EvaluationYunhao Tang, Tadashi Kozuno, Mark Rowland, Rémi Munos et al.NeurIPS 2021 · 9 citations
- Biased Gradient Estimate with Drastic Variance Reduction for Meta Reinforcement LearningYunhao TangICML 2022 · 7 citations
- Distributional Meta-Gradient Reinforcement LearningHaiyan Yin, Shuicheng Yan, Zhongwen XuICLR 2023
- On the Convergence Theory of Debiased Model-Agnostic Meta-Reinforcement LearningAlireza Fallah, Kristian Georgiev, Aryan Mokhtari, Asuman E. OzdaglarNeurIPS 2021 · 31 citations
- Statistically Efficient Off-Policy Policy GradientsNathan Kallus, Masatoshi UeharaICML 2020 · 43 citations
