Lune

ICML2022顶会

Biased Gradient Estimate with Drastic Variance Reduction for Meta Reinforcement Learning

Yunhao Tang

2022年份
7被引次数
5顶会引用

摘要

Despite the empirical success of meta reinforcement learning (meta-RL), there are still a number poorly-understood discrepancies between theory and practice. Critically, biased gradient estimates are almost always implemented in practice, whereas prior theory on meta-RL only establishes convergence under unbiased gradient estimates. In this work, we investigate such a discrepancy. In particular, (1) We show that unbiased gradient estimates have variance Θ(N)\Theta(N) which linearly depends on the sample size NN of the inner loop updates; (2) We propose linearized score function (LSF) gradient estimates, which have bias O(1/N)\mathcal{O}(1/\sqrt{N}) and variance O(1/N)\mathcal{O}(1/N); (3) We show that most empirical prior work in fact implements variants of the LSF gradient estimates. This implies that practical algorithms"accidentally"introduce bias to achieve better performance; (4) We establish theoretical guarantees for the LSF gradient estimates in meta-RL regarding its convergence to stationary points, showing better dependency on NN than prior work when NN is large.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper5

问问它们各自怎么用它

它引用的顶会 Paper4

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖