Lune

ICLR2020顶会

Sample Efficient Policy Gradient Methods with Recursive Variance Reduction

Pan Xu, Felicia Gao, Quanquan Gu

2020年份
99被引次数
40顶会引用

摘要

Improving the sample efficiency in reinforcement learning has been a long-standing research problem. In this work, we aim to reduce the sample complexity of existing policy gradient methods. We propose a novel policy gradient algorithm called SRVR-PG, which only requires O(1/ϵ3/2)O(1/\epsilon^{3/2}) episodes to find an ϵ\epsilon-approximate stationary point of the nonconcave performance function J(θ)J(\boldsymbol{\theta}) (i.e., θ\boldsymbol{\theta} such that ∥∇J(θ)∥22≤ϵ\|\nabla J(\boldsymbol{\theta})\|_2^2\leq\epsilon). This sample complexity improves the existing result O(1/ϵ5/3)O(1/\epsilon^{5/3}) for stochastic variance reduced policy gradient algorithms by a factor of O(1/ϵ1/6)O(1/\epsilon^{1/6}). In addition, we also propose a variant of SRVR-PG with parameter exploration, which explores the initial policy parameter from a prior probability distribution. We conduct numerical experiments on classic control problems in reinforcement learning to validate the performance of our proposed algorithms.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper40

问问它们各自怎么用它

它引用的顶会 Paper2

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖