Unifying Gradient Estimators for Meta-Reinforcement Learning via Off-Policy Evaluation
Yunhao Tang, Tadashi Kozuno, Mark Rowland, Rémi Munos, Michal Valko
Abstract
Model-agnostic meta-reinforcement learning requires estimating the Hessian matrix of value functions. This is challenging from an implementation perspective, as repeatedly differentiating policy gradient estimates may lead to biased Hessian estimates. In this work, we provide a unifying framework for estimating higherorder derivatives of value functions, based on off-policy evaluation. Our framework interprets a number of prior approaches as special cases and elucidates the bias and variance trade-off of Hessian estimates. This framework also opens the door to a new family of estimates, which can be easily implemented with auto-differentiation libraries, and lead to performance gains in practice. We open source the code to reproduce our results 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- A Theoretical Understanding of Gradient Bias in Meta-Reinforcement LearningBo Liu, Xidong Feng, Jie Ren, Luo Mai et al.NeurIPS 2022 · 16 citations
- Biased Gradient Estimate with Drastic Variance Reduction for Meta Reinforcement LearningYunhao TangICML 2022 · 7 citations
- Optimizing Language Models for Inference Time Objectives using Reinforcement LearningYunhao Tang, Kunhao Zheng, Gabriel Synnaeve, Rémi MunosICML 2025
Builds on5
- Discovering Reinforcement Learning AlgorithmsJunhyuk Oh, Matteo Hessel, Wojciech M. Czarnecki, Zhongwen Xu et al.NeurIPS 2020 · 154 citations
- A Self-Tuning Actor-Critic AlgorithmTom Zahavy, Zhongwen Xu, Vivek Veeriah, Matteo Hessel et al.NeurIPS 2020 · 106 citations
- Meta-Gradient Reinforcement Learning with an Objective Discovered OnlineZhongwen Xu, Hado Philip van Hasselt, Matteo Hessel, Junhyuk Oh et al.NeurIPS 2020 · 90 citations
- From Importance Sampling to Doubly Robust Policy GradientJiawei Huang, Nan JiangICML 2020 · 26 citations
- Taylor Expansion Policy OptimizationYunhao Tang, Michal Valko, Rémi MunosICML 2020 · 16 citations
Related papers
- ES-MAML: Simple Hessian-Free Meta LearningXingyou Song, Wenbo Gao, Yuxiang Yang, Krzysztof Choromanski et al.ICLR 2020 · 128 citations
- Statistically Efficient Off-Policy Policy GradientsNathan Kallus, Masatoshi UeharaICML 2020 · 43 citations
- On the Second-Order Convergence of Biased Policy Gradient AlgorithmsSiqiao Mu, Diego KlabjanICML 2024 · 4 citations
- Revisiting Scalable Hessian Diagonal Approximations for Applications in Reinforcement LearningMohamed Elsayed, Homayoon Farrahi, Felix Dangel, A. Rupam MahmoodICML 2024 · 7 citations
- Revisiting Zeroth-Order Hessian Approximation: A Single-Step Policy Optimization LensJunbin Qiu, Zhaowei Hong, Renzhe Xu, Yao ShuICML 2026
