Biased Gradient Estimate with Drastic Variance Reduction for Meta Reinforcement Learning
Yunhao Tang
Abstract
Despite the empirical success of meta reinforcement learning (meta-RL), there are still a number poorly-understood discrepancies between theory and practice. Critically, biased gradient estimates are almost always implemented in practice, whereas prior theory on meta-RL only establishes convergence under unbiased gradient estimates. In this work, we investigate such a discrepancy. In particular, (1) We show that unbiased gradient estimates have variance which linearly depends on the sample size of the inner loop updates; (2) We propose linearized score function (LSF) gradient estimates, which have bias and variance ; (3) We show that most empirical prior work in fact implements variants of the LSF gradient estimates. This implies that practical algorithms"accidentally"introduce bias to achieve better performance; (4) We establish theoretical guarantees for the LSF gradient estimates in meta-RL regarding its convergence to stationary points, showing better dependency on than prior work when is large.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 070c3d14-86b6-4dfa-a227-5e1d52a5318fCited by top-tier papers5
- MetaRLEC: Meta-Reinforcement Learning for Discovery of Brain Effective ConnectivityZuozhen Zhang, Junzhong Ji, Jinduo LiuAAAI 2024 · 12 citations
- Meta-Reinforcement Learning with Universal Policy Adaptation: Provable Near-Optimality under All-task Optimum ComparatorSiyuan Xu, Minghui ZhuNeurIPS 2024 · 8 citations
- MetaCARD: Meta-Reinforcement Learning with Task Uncertainty Feedback via Decoupled Context-Aware Reward and Dynamics ComponentsMin Wang, Xin Li, Leiji Zhang, Mingzhong WangAAAI 2024 · 6 citations
- MetaTrader: Learning to Generalize RL Trading Policies Beyond Offline DataHaochen Yuan, Minting Pan, Yunbo Wang, Siyu Gao et al.AAAI 2026
- Optimizing Language Models for Inference Time Objectives using Reinforcement LearningYunhao Tang, Kunhao Zheng, Gabriel Synnaeve, Rémi MunosICML 2025
Builds on4
- Discovering Reinforcement Learning AlgorithmsJunhyuk Oh, Matteo Hessel, Wojciech M. Czarnecki, Zhongwen Xu et al.NeurIPS 2020 · 154 citations
- Meta-Gradient Reinforcement Learning with an Objective Discovered OnlineZhongwen Xu, Hado Philip van Hasselt, Matteo Hessel, Junhyuk Oh et al.NeurIPS 2020 · 90 citations
- Taylor Expansion Policy OptimizationYunhao Tang, Michal Valko, Rémi MunosICML 2020 · 16 citations
- Unifying Gradient Estimators for Meta-Reinforcement Learning via Off-Policy EvaluationYunhao Tang, Tadashi Kozuno, Mark Rowland, Rémi Munos et al.NeurIPS 2021 · 9 citations
Related papers
- On the Second-Order Convergence of Biased Policy Gradient AlgorithmsSiqiao Mu, Diego KlabjanICML 2024 · 4 citations
- On the Convergence Theory of Debiased Model-Agnostic Meta-Reinforcement LearningAlireza Fallah, Kristian Georgiev, Aryan Mokhtari, Asuman E. OzdaglarNeurIPS 2021 · 31 citations
- A Theoretical Understanding of Gradient Bias in Meta-Reinforcement LearningBo Liu, Xidong Feng, Jie Ren, Luo Mai et al.NeurIPS 2022 · 16 citations
- Statistically Efficient Off-Policy Policy GradientsNathan Kallus, Masatoshi UeharaICML 2020 · 43 citations
- Theoretical Guarantees of Fictitious Discount Algorithms for Episodic Reinforcement Learning and Global Convergence of Policy Gradient MethodsXin Guo, Anran Hu, Junzi ZhangAAAI 2022 · 10 citations
