Lune

NeurIPS2021顶会

Agnostic Reinforcement Learning with Low-Rank MDPs and Rich Observations

Ayush Sekhari, Christoph Dann, Mehryar Mohri, Yishay Mansour, Karthik Sridharan

2021年份
15被引次数
6顶会引用

摘要

There have been many recent advances on provably efficient Reinforcement Learning (RL) in problems with rich observation spaces. However, all these works share a strong realizability assumption about the optimal value function of the true MDP. Such realizability assumptions are often too strong to hold in practice. In this work, we consider the more realistic setting of agnostic RL with rich observation spaces and a fixed class of policies Π\Pi that may not contain any near-optimal policy. We provide an algorithm for this setting whose error is bounded in terms of the rank dd of the underlying MDP. Specifically, our algorithm enjoys a sample complexity bound of O~((H4dK3dlog⁡∣Π∣)/ϵ2)\widetilde{O}\left((H^{4d} K^{3d} \log |\Pi|)/\epsilon^2\right) where HH is the length of episodes, KK is the number of actions and ϵ>0\epsilon>0 is the desired sub-optimality. We also provide a nearly matching lower bound for this agnostic setting that shows that the exponential dependence on rank is unavoidable, without further assumptions.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper6

问问它们各自怎么用它

它引用的顶会 Paper13

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖