Bayesian Bellman Operators
Mattie Fellows, Kristian Hartikainen, Shimon Whiteson
Abstract
We introduce a novel perspective on Bayesian reinforcement learning (RL); whereas existing approaches infer a posterior over the transition distribution or Q-function, we characterise the uncertainty in the Bellman operator. Our Bayesian Bellman operator (BBO) framework is motivated by the insight that when bootstrapping is introduced, model-free approaches actually infer a posterior over Bellman operators, not value functions. In this paper, we use BBO to provide a rigorous theoretical analysis of model-free Bayesian RL to better understand its relationship to established frequentist RL methodologies. We prove that Bayesian solutions are consistent with frequentist RL solutions, even when approximate inference is used, and derive conditions for which convergence properties hold. Empirically, we demonstrate that algorithms derived from the BBO framework have sophisticated deep exploration properties that enable them to solve continuous control tasks at which state-of-the-art regularised actor-critic algorithms fail catastrophically.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- An Adaptive Deep RL Method for Non-Stationary Environments with Piecewise Stable ContextXiaoyu Chen, Xiangming Zhu, Yufeng Zheng, Pushi Zhang et al.NeurIPS 2022 · 24 citations
- Uncertainty-based Offline Variational Bayesian Reinforcement Learning for Robustness under Diverse Data CorruptionsRui Yang, Jie Wang, Guoping Wu, Bin LiNeurIPS 2024 · 11 citations
- Diverse Projection Ensembles for Distributional Reinforcement LearningMoritz Akiya Zanger, Wendelin Boehmer, Matthijs T. J. SpaanICLR 2024 · 9 citations
- Bayesian Exploration NetworksMattie Fellows, Brandon Kaplowitz, Christian Schröder de Witt, Shimon WhitesonICML 2024 · 4 citations
- Epistemic Bellman OperatorsPascal R. van der Vaart, Matthijs T. J. Spaan, Neil Yorke-SmithAAAI 2025 · 2 citations
Builds on3
- VariBAD: A Very Good Method for Bayes-Adaptive Deep RL via Meta-LearningLuisa M. Zintgraf, Kyriacos Shiarlis, Maximilian Igl, Sebastian Schulze et al.ICLR 2020 · 315 citations
- Conservative Uncertainty Estimation By Fitting Prior NetworksKamil Ciosek, Vincent Fortuin, Ryota Tomioka, Katja Hofmann et al.ICLR 2020 · 65 citations
- Geometric Insights into the Convergence of Nonlinear TD LearningDavid Brandfonbrener, Joan BrunaICLR 2020 · 18 citations
Related papers
- Principled Exploration via Optimistic Bootstrapping and Backward InductionChenjia Bai, Lingxiao Wang, Lei Han, Jianye Hao et al.ICML 2021 · 46 citations
- ReLU to the Rescue: Improve Your On-Policy Actor-Critic with Positive AdvantagesAndrew Jesson, Chris Lu, Gunshi Gupta, Nicolas Beltran-Velez et al.ICML 2024 · 11 citations
- Parameterized Projected Bellman OperatorThéo Vincent, Alberto Maria Metelli, Boris Belousov, Jan Peters et al.AAAI 2024 · 6 citations
- Sample Efficient Deep Reinforcement Learning via Uncertainty EstimationVincent Mai, Kaustubh Mani, Liam PaullICLR 2022 · 53 citations
- SUNRISE: A Simple Unified Framework for Ensemble Learning in Deep Reinforcement LearningKimin Lee, Michael Laskin, Aravind Srinivas, Pieter AbbeelICML 2021 · 239 citations
