Lune

NeurIPS2020顶会

Sample Efficient Reinforcement Learning via Low-Rank Matrix Estimation

Devavrat Shah, Dogyoon Song, Zhi Xu, Yuzhe Yang

2020年份
35被引次数
11顶会引用

摘要

We consider the question of learning QQ-function in a sample efficient manner for reinforcement learning with continuous state and action spaces under a generative model. If QQ-function is Lipschitz continuous, then the minimal sample complexity for estimating ϵ\epsilon-optimal QQ-function is known to scale as Ω(1ϵd1+d2+2){\Omega}(\frac{1}{\epsilon^{d_1+d_2 +2}}) per classical non-parametric learning theory, where d1d_1 and d2d_2 denote the dimensions of the state and action spaces respectively. The QQ-function, when viewed as a kernel, induces a Hilbert-Schmidt operator and hence possesses square-summable spectrum. This motivates us to consider a parametric class of QQ-functions parameterized by its "rank" rr, which contains all Lipschitz QQ-functions as r→∞r \to \infty. As our key contribution, we develop a simple, iterative learning algorithm that finds ϵ\epsilon-optimal QQ-function with sample complexity of O~(1ϵmax⁡(d1,d2)+2)\widetilde{O}(\frac{1}{\epsilon^{\max(d_1, d_2)+2}}) when the optimal QQ-function has low rank rr and the discounting factor γ\gamma is below a certain threshold. Thus, this provides an exponential improvement in sample complexity. To enable our result, we develop a novel Matrix Estimation algorithm that faithfully estimates an unknown low-rank matrix in the ℓ∞\ell_\infty sense even in the presence of arbitrary bounded noise, which might be of interest in its own right. Empirical results on several stochastic control tasks confirm the efficacy of our "low-rank" algorithms.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper11

问问它们各自怎么用它

它引用的顶会 Paper2

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖