From Dirichlet to Rubin: Optimistic Exploration in RL without Bonuses
Daniil Tiapkin, Denis Belomestny, Eric Moulines, Alexey Naumov, Sergey Samsonov, Yunhao Tang, Michal Valko, Pierre Ménard
摘要
We propose the Bayes-UCBVI algorithm for reinforcement learning in tabular, stage-dependent, episodic Markov decision process: a natural extension of the Bayes-UCB algorithm by Kaufmann et al. (2012) for multi-armed bandits. Our method uses the quantile of a Q-value function posterior as upper confidence bound on the optimal Q-value function. For Bayes-UCBVI, we prove a regret bound of order O( √ H 3 SAT ) where H is the length of one episode, S is the number of states, A the number of actions, T the number of episodes, that matches the lower-bound of Ω( √ H 3 SAT ) up to poly-log terms in H, S, A, T for a large enough T . To the best of our knowledge, this is the first algorithm that obtains an optimal dependence on the horizon H (and S) without the need of an involved Bernstein-like bonus or noise. Crucial to our analysis is a new fine-grained anticoncentration bound for a weighted Dirichlet sum that can be of independent interest. We then explain how Bayes-UCBVI can be easily extended beyond the tabular setting, exhibiting a strong link between our algorithm and Bayesian bootstrap (Rubin, 1981). 1 We translate all the bounds to the stage-dependent setting by multiplying by √ H the regret bounds in the stage-independent setting. 2 In the O(•) notation we ignore terms poly-log in H, S, A, T . 3 Or simple linearly parameterized settings.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Provable and Practical: Efficient Exploration in Reinforcement Learning via Langevin Monte CarloHaque Ishfaq, Qingfeng Lan, Pan Xu, A. Rupam Mahmood 等ICLR 2024 · 被引用 33 次
- Optimistic Posterior Sampling for Reinforcement Learning with Few Samples and Tight GuaranteesDaniil Tiapkin, Denis Belomestny, Daniele Calandriello, Eric Moulines 等NeurIPS 2022 · 被引用 16 次
- Q-Star Meets Scalable Posterior Sampling: Bridging Theory and Practice via HyperAgentYingru Li, Jiawei Xu, Lei Han, Zhi-Quan LuoICML 2024 · 被引用 8 次
- Model-free Posterior Sampling via Learning Rate RandomizationDaniil Tiapkin, Denis Belomestny, Daniele Calandriello, Eric Moulines 等NeurIPS 2023 · 被引用 8 次
- Provably and Practically Efficient Adversarial Imitation Learning with General Function ApproximationTian Xu, Zhilong Zhang, Ruishuo Chen, Yihao Sun 等NeurIPS 2024 · 被引用 8 次
它引用的顶会 Paper3
- Never Give Up: Learning Directed Exploration StrategiesAdrià Puigdomènech Badia, Pablo Sprechmann, Alex Vitvitskyi, Zhaohan Daniel Guo 等ICLR 2020 · 被引用 349 次
- Principled Exploration via Optimistic Bootstrapping and Backward InductionChenjia Bai, Lingxiao Wang, Lei Han, Jianye Hao 等ICML 2021 · 被引用 46 次
- Planning in Markov Decision Processes with Gap-Dependent Sample ComplexityAnders Jonsson, Emilie Kaufmann, Pierre Ménard, Omar Darwiche Domingues 等NeurIPS 2020 · 被引用 46 次
相关 Paper
- UCB Momentum Q-learning: Correcting the bias without forgettingPierre Ménard, Omar Darwiche Domingues, Xuedong Shang, Michal ValkoICML 2021 · 被引用 53 次
- Nearly Minimax Optimal Reinforcement Learning for Discounted MDPsJiafan He, Dongruo Zhou, Quanquan GuNeurIPS 2021 · 被引用 53 次
- Q-Learning with Fine-Grained Gap-Dependent RegretHaochen Zhang, Zhong Zheng, Lingzhou XueICLR 2026 · 被引用 3 次
- Almost Optimal Model-Free Reinforcement Learningvia Reference-Advantage DecompositionZihan Zhang, Yuan Zhou, Xiangyang JiNeurIPS 2020 · 被引用 183 次
- Improved Bayesian Regret Bounds for Thompson Sampling in Reinforcement LearningAhmadreza Moradipari, Mohammad Pedramfar, Modjtaba Shokrian Zini, Vaneet AggarwalNeurIPS 2023 · 被引用 8 次
