Biased Gradient Estimate with Drastic Variance Reduction for Meta Reinforcement Learning
Yunhao Tang
摘要
Despite the empirical success of meta reinforcement learning (meta-RL), there are still a number poorly-understood discrepancies between theory and practice. Critically, biased gradient estimates are almost always implemented in practice, whereas prior theory on meta-RL only establishes convergence under unbiased gradient estimates. In this work, we investigate such a discrepancy. In particular, (1) We show that unbiased gradient estimates have variance which linearly depends on the sample size of the inner loop updates; (2) We propose linearized score function (LSF) gradient estimates, which have bias and variance ; (3) We show that most empirical prior work in fact implements variants of the LSF gradient estimates. This implies that practical algorithms"accidentally"introduce bias to achieve better performance; (4) We establish theoretical guarantees for the LSF gradient estimates in meta-RL regarding its convergence to stationary points, showing better dependency on than prior work when is large.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- MetaRLEC: Meta-Reinforcement Learning for Discovery of Brain Effective ConnectivityZuozhen Zhang, Junzhong Ji, Jinduo LiuAAAI 2024 · 被引用 12 次
- Meta-Reinforcement Learning with Universal Policy Adaptation: Provable Near-Optimality under All-task Optimum ComparatorSiyuan Xu, Minghui ZhuNeurIPS 2024 · 被引用 8 次
- MetaCARD: Meta-Reinforcement Learning with Task Uncertainty Feedback via Decoupled Context-Aware Reward and Dynamics ComponentsMin Wang, Xin Li, Leiji Zhang, Mingzhong WangAAAI 2024 · 被引用 6 次
- MetaTrader: Learning to Generalize RL Trading Policies Beyond Offline DataHaochen Yuan, Minting Pan, Yunbo Wang, Siyu Gao 等AAAI 2026
- Optimizing Language Models for Inference Time Objectives using Reinforcement LearningYunhao Tang, Kunhao Zheng, Gabriel Synnaeve, Rémi MunosICML 2025
它引用的顶会 Paper4
- Discovering Reinforcement Learning AlgorithmsJunhyuk Oh, Matteo Hessel, Wojciech M. Czarnecki, Zhongwen Xu 等NeurIPS 2020 · 被引用 154 次
- Meta-Gradient Reinforcement Learning with an Objective Discovered OnlineZhongwen Xu, Hado Philip van Hasselt, Matteo Hessel, Junhyuk Oh 等NeurIPS 2020 · 被引用 90 次
- Taylor Expansion Policy OptimizationYunhao Tang, Michal Valko, Rémi MunosICML 2020 · 被引用 16 次
- Unifying Gradient Estimators for Meta-Reinforcement Learning via Off-Policy EvaluationYunhao Tang, Tadashi Kozuno, Mark Rowland, Rémi Munos 等NeurIPS 2021 · 被引用 9 次
相关 Paper
- On the Second-Order Convergence of Biased Policy Gradient AlgorithmsSiqiao Mu, Diego KlabjanICML 2024 · 被引用 4 次
- On the Convergence Theory of Debiased Model-Agnostic Meta-Reinforcement LearningAlireza Fallah, Kristian Georgiev, Aryan Mokhtari, Asuman E. OzdaglarNeurIPS 2021 · 被引用 31 次
- A Theoretical Understanding of Gradient Bias in Meta-Reinforcement LearningBo Liu, Xidong Feng, Jie Ren, Luo Mai 等NeurIPS 2022 · 被引用 16 次
- Statistically Efficient Off-Policy Policy GradientsNathan Kallus, Masatoshi UeharaICML 2020 · 被引用 43 次
- Theoretical Guarantees of Fictitious Discount Algorithms for Episodic Reinforcement Learning and Global Convergence of Policy Gradient MethodsXin Guo, Anran Hu, Junzi ZhangAAAI 2022 · 被引用 10 次
