Lune

ICLR2023顶会

Towards Minimax Optimal Reward-free Reinforcement Learning in Linear MDPs

Pihe Hu, Yu Chen, Longbo Huang

出版方
2023年份
5顶会引用

摘要

We study reward-free reinforcement learning with linear function approximation for episodic Markov decision processes (MDPs). In this setting, an agent first interacts with the environment without accessing the reward function in the exploration phase. In the subsequent planning phase, it is given a reward function and asked to output an ϵ\epsilon-optimal policy. We propose a novel algorithm LSVI-RFE under the linear MDP setting, where the transition probability and reward functions are linear in a feature mapping. We prove an O~(H4d2/ϵ2)\widetilde{O}(H^{4} d^{2}/\epsilon^2) sample complexity upper bound for LSVI-RFE, where HH is the episode length and dd is the feature dimension. We also establish a sample complexity lower bound of Ω(H3d2/ϵ2)\Omega(H^{3} d^{2}/\epsilon^2). To the best of our knowledge, LSVI-RFE is the first computationally efficient algorithm that achieves the minimax optimal sample complexity in linear MDP settings up to an HH and logarithmic factors. Our LSVI-RFE algorithm is based on a novel variance-aware exploration mechanism to avoid overly-conservative exploration in prior works. Our sharp bound relies on the decoupling of UCB bonuses during two phases, and a Bernstein-type self-normalized bound, which remove the extra dependency of sample complexity on HH and dd, respectively.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper5

问问它们各自怎么用它

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖