Lune

ICML2020Top-tier venue

Optimistic Policy Optimization with Bandit Feedback

Lior Shani, Yonathan Efroni, Aviv Rosenberg, Shie Mannor

2020Year
100Citations
61Top-tier citations

Abstract

Policy optimization methods are one of the most widely used classes of Reinforcement Learning (RL) algorithms. Yet, so far, such methods have been mostly analyzed from an optimization perspective, without addressing the problem of exploration, or by making strong assumptions on the interaction with the environment. In this paper we consider model-based RL in the tabular finite-horizon MDP setting with unknown transitions and bandit feedback. For this setting, we propose an optimistic policy optimization algorithm for which we establish Õ( √ S 2 AH 4 K) regret for stochastic rewards. Furthermore, we prove Õ( √ S 2 AH 4 K 2/3 ) regret for adversarial rewards. Interestingly, this result matches previous bounds derived for the bandit feedback case, yet with known transitions. To the best of our knowledge, the two results are the first sub-linear regret bounds obtained for policy optimization algorithms with unknown transitions and bandit feedback. * Equal contribution, ordering decided by a coin toss.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext d5bad5ad-a06c-4345-87ee-730dc740906f

Cited by top-tier papers61

Ask how each one uses it

Builds on4

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines