Lune

NeurIPS2021Top-tier venue

Nearly Minimax Optimal Reinforcement Learning for Discounted MDPs

Jiafan He, Dongruo Zhou, Quanquan Gu

2021Year
53Citations
15Top-tier citations

Abstract

We study the reinforcement learning problem for discounted Markov Decision Processes (MDPs) under the tabular setting. We propose a model-based algorithm named UCBVI-γ\gamma, which is based on the optimism in the face of uncertainty principle and the Bernstein-type bonus. We show that UCBVI-γ\gamma achieves an O~(SAT/(1−γ)1.5)\tilde{O}\big({\sqrt{SAT}}/{(1-\gamma)^{1.5}}\big) regret, where SS is the number of states, AA is the number of actions, γ\gamma is the discount factor and TT is the number of steps. In addition, we construct a class of hard MDPs and show that for any algorithm, the expected regret is at least Ω~(SAT/(1−γ)1.5)\tilde{\Omega}\big({\sqrt{SAT}}/{(1-\gamma)^{1.5}}\big). Our upper bound matches the minimax lower bound up to logarithmic factors, which suggests that UCBVI-γ\gamma is nearly minimax optimal for discounted MDPs.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 22ed8be6-69fb-47e7-afbd-175eb37983ad

Cited by top-tier papers15

Ask how each one uses it

Builds on6

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines