Lune

ICML2022Top-tier venue

Biased Gradient Estimate with Drastic Variance Reduction for Meta Reinforcement Learning

Yunhao Tang

2022Year
7Citations
5Top-tier citations

Abstract

Despite the empirical success of meta reinforcement learning (meta-RL), there are still a number poorly-understood discrepancies between theory and practice. Critically, biased gradient estimates are almost always implemented in practice, whereas prior theory on meta-RL only establishes convergence under unbiased gradient estimates. In this work, we investigate such a discrepancy. In particular, (1) We show that unbiased gradient estimates have variance Θ(N)\Theta(N) which linearly depends on the sample size NN of the inner loop updates; (2) We propose linearized score function (LSF) gradient estimates, which have bias O(1/N)\mathcal{O}(1/\sqrt{N}) and variance O(1/N)\mathcal{O}(1/N); (3) We show that most empirical prior work in fact implements variants of the LSF gradient estimates. This implies that practical algorithms"accidentally"introduce bias to achieve better performance; (4) We establish theoretical guarantees for the LSF gradient estimates in meta-RL regarding its convergence to stationary points, showing better dependency on NN than prior work when NN is large.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 070c3d14-86b6-4dfa-a227-5e1d52a5318f

Cited by top-tier papers5

Ask how each one uses it

Builds on4

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines