From Dirichlet to Rubin: Optimistic Exploration in RL without Bonuses
Daniil Tiapkin, Denis Belomestny, Eric Moulines, Alexey Naumov, Sergey Samsonov, Yunhao Tang, Michal Valko, Pierre Ménard
Abstract
We propose the Bayes-UCBVI algorithm for reinforcement learning in tabular, stage-dependent, episodic Markov decision process: a natural extension of the Bayes-UCB algorithm by Kaufmann et al. (2012) for multi-armed bandits. Our method uses the quantile of a Q-value function posterior as upper confidence bound on the optimal Q-value function. For Bayes-UCBVI, we prove a regret bound of order O( √ H 3 SAT ) where H is the length of one episode, S is the number of states, A the number of actions, T the number of episodes, that matches the lower-bound of Ω( √ H 3 SAT ) up to poly-log terms in H, S, A, T for a large enough T . To the best of our knowledge, this is the first algorithm that obtains an optimal dependence on the horizon H (and S) without the need of an involved Bernstein-like bonus or noise. Crucial to our analysis is a new fine-grained anticoncentration bound for a weighted Dirichlet sum that can be of independent interest. We then explain how Bayes-UCBVI can be easily extended beyond the tabular setting, exhibiting a strong link between our algorithm and Bayesian bootstrap (Rubin, 1981). 1 We translate all the bounds to the stage-dependent setting by multiplying by √ H the regret bounds in the stage-independent setting. 2 In the O(•) notation we ignore terms poly-log in H, S, A, T . 3 Or simple linearly parameterized settings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- Provable and Practical: Efficient Exploration in Reinforcement Learning via Langevin Monte CarloHaque Ishfaq, Qingfeng Lan, Pan Xu, A. Rupam Mahmood et al.ICLR 2024 · 33 citations
- Optimistic Posterior Sampling for Reinforcement Learning with Few Samples and Tight GuaranteesDaniil Tiapkin, Denis Belomestny, Daniele Calandriello, Eric Moulines et al.NeurIPS 2022 · 16 citations
- Q-Star Meets Scalable Posterior Sampling: Bridging Theory and Practice via HyperAgentYingru Li, Jiawei Xu, Lei Han, Zhi-Quan LuoICML 2024 · 8 citations
- Model-free Posterior Sampling via Learning Rate RandomizationDaniil Tiapkin, Denis Belomestny, Daniele Calandriello, Eric Moulines et al.NeurIPS 2023 · 8 citations
- Provably and Practically Efficient Adversarial Imitation Learning with General Function ApproximationTian Xu, Zhilong Zhang, Ruishuo Chen, Yihao Sun et al.NeurIPS 2024 · 8 citations
Builds on3
- Never Give Up: Learning Directed Exploration StrategiesAdrià Puigdomènech Badia, Pablo Sprechmann, Alex Vitvitskyi, Zhaohan Daniel Guo et al.ICLR 2020 · 349 citations
- Principled Exploration via Optimistic Bootstrapping and Backward InductionChenjia Bai, Lingxiao Wang, Lei Han, Jianye Hao et al.ICML 2021 · 46 citations
- Planning in Markov Decision Processes with Gap-Dependent Sample ComplexityAnders Jonsson, Emilie Kaufmann, Pierre Ménard, Omar Darwiche Domingues et al.NeurIPS 2020 · 46 citations
Related papers
- UCB Momentum Q-learning: Correcting the bias without forgettingPierre Ménard, Omar Darwiche Domingues, Xuedong Shang, Michal ValkoICML 2021 · 53 citations
- Nearly Minimax Optimal Reinforcement Learning for Discounted MDPsJiafan He, Dongruo Zhou, Quanquan GuNeurIPS 2021 · 53 citations
- Q-Learning with Fine-Grained Gap-Dependent RegretHaochen Zhang, Zhong Zheng, Lingzhou XueICLR 2026 · 3 citations
- Almost Optimal Model-Free Reinforcement Learningvia Reference-Advantage DecompositionZihan Zhang, Yuan Zhou, Xiangyang JiNeurIPS 2020 · 183 citations
- Improved Bayesian Regret Bounds for Thompson Sampling in Reinforcement LearningAhmadreza Moradipari, Mohammad Pedramfar, Modjtaba Shokrian Zini, Vaneet AggarwalNeurIPS 2023 · 8 citations
