Lune

NeurIPS2021Top-tier venue

Policy Optimization in Adversarial MDPs: Improved Exploration via Dilated Bonuses

Haipeng Luo, Chen-Yu Wei, Chung-Wei Lee

2021Year
59Citations
41Top-tier citations

Abstract

Policy optimization is a widely-used method in reinforcement learning. Due to its local-search nature, however, theoretical guarantees on global optimality often rely on extra assumptions on the Markov Decision Processes (MDPs) that bypass the challenge of global exploration. To eliminate the need of such assumptions, in this work, we develop a general solution that adds dilated bonuses to the policy update to facilitate global exploration. To showcase the power and generality of this technique, we apply it to several episodic MDP settings with adversarial losses and bandit feedback, improving and generalizing the state-of-the-art. Specifically, in the tabular case, we obtain O~(T)\widetilde{\mathcal{O}}(\sqrt{T}) regret where TT is the number of episodes, improving the O~(T2/3)\widetilde{\mathcal{O}}({T}^{2/3}) regret bound by Shani et al. (2020). When the number of states is infinite, under the assumption that the state-action values are linear in some low-dimensional features, we obtain O~(T2/3)\widetilde{\mathcal{O}}({T}^{2/3}) regret with the help of a simulator, matching the result of Neu and Olkhovskaya (2020) while importantly removing the need of an exploratory policy that their algorithm requires. When a simulator is unavailable, we further consider a linear MDP setting and obtain O~(T14/15)\widetilde{\mathcal{O}}({T}^{14/15}) regret, which is the first result for linear MDPs with adversarial losses and bandit feedback.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 37858e3a-a535-49e7-b709-3a3fb293e9fd

Cited by top-tier papers41

Ask how each one uses it

Builds on12

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines