Lune

ICLR2020顶会

Q-learning with UCB Exploration is Sample Efficient for Infinite-Horizon MDP

Yuanhao Wang, Kefan Dong, Xiaoyu Chen, Liwei Wang

2020年份
107被引次数
46顶会引用

摘要

A fundamental question in reinforcement learning is whether model-free algorithms are sample efficient. Recently, Jin et al. proposed a Q-learning algorithm with UCB exploration policy, and proved it has nearly optimal regret bound for finite-horizon episodic MDP. In this paper, we adapt Q-learning with UCB-exploration bonus to infinite-horizon MDP with discounted rewards without accessing a generative model. We show that the sample complexity of exploration of our algorithm is bounded by O~(SAϵ2(1−γ)7)\tilde{O}({\frac{SA}{\epsilon^2(1-\gamma)^7}}). This improves the previously best known result of O~(SAϵ4(1−γ)8)\tilde{O}({\frac{SA}{\epsilon^4(1-\gamma)^8}}) in this setting achieved by delayed Q-learning , and matches the lower bound in terms of ϵ\epsilon as well as SS and AA except for logarithmic factors.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper46

问问它们各自怎么用它

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖