Lune

ICML2022顶会

Reward-Free RL is No Harder Than Reward-Aware RL in Linear Markov Decision Processes

Andrew J. Wagenmaker, Yifang Chen, Max Simchowitz, Simon S. Du, Kevin Jamieson

2022年份
61被引次数
34顶会引用

摘要

Reward-free reinforcement learning (RL) considers the setting where the agent does not have access to a reward function during exploration, but must propose a near-optimal policy for an arbitrary reward function revealed only after exploring. In the the tabular setting, it is well known that this is a more difficult problem than reward-aware (PAC) RL -- where the agent has access to the reward function during exploration -- with optimal sample complexities in the two settings differing by a factor of ∣S∣|\mathcal{S}|, the size of the state space. We show that this separation does not exist in the setting of linear MDPs. We first develop a computationally efficient algorithm for reward-free RL in a dd-dimensional linear MDP with sample complexity scaling as O~(d2H5/ϵ2)\widetilde{\mathcal{O}}(d^2 H^5/\epsilon^2). We then show a lower bound with matching dimension-dependence of Ω(d2H2/ϵ2)\Omega(d^2 H^2/\epsilon^2), which holds for the reward-aware RL setting. To our knowledge, our approach is the first computationally efficient algorithm to achieve optimal dd dependence in linear MDPs, even in the single-reward PAC setting. Our algorithm relies on a novel procedure which efficiently traverses a linear MDP, collecting samples in any given feature direction'', and enjoys a sample complexity scaling optimally in the (linear MDP equivalent of the) maximal state visitation probability. We show that this exploration procedure can also be applied to solve the problem of obtaining well-conditioned'' covariates in linear MDPs.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper34

问问它们各自怎么用它

它引用的顶会 Paper15

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖