Bayesian Bellman Operators
Mattie Fellows, Kristian Hartikainen, Shimon Whiteson
摘要
We introduce a novel perspective on Bayesian reinforcement learning (RL); whereas existing approaches infer a posterior over the transition distribution or Q-function, we characterise the uncertainty in the Bellman operator. Our Bayesian Bellman operator (BBO) framework is motivated by the insight that when bootstrapping is introduced, model-free approaches actually infer a posterior over Bellman operators, not value functions. In this paper, we use BBO to provide a rigorous theoretical analysis of model-free Bayesian RL to better understand its relationship to established frequentist RL methodologies. We prove that Bayesian solutions are consistent with frequentist RL solutions, even when approximate inference is used, and derive conditions for which convergence properties hold. Empirically, we demonstrate that algorithms derived from the BBO framework have sophisticated deep exploration properties that enable them to solve continuous control tasks at which state-of-the-art regularised actor-critic algorithms fail catastrophically.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- An Adaptive Deep RL Method for Non-Stationary Environments with Piecewise Stable ContextXiaoyu Chen, Xiangming Zhu, Yufeng Zheng, Pushi Zhang 等NeurIPS 2022 · 被引用 24 次
- Uncertainty-based Offline Variational Bayesian Reinforcement Learning for Robustness under Diverse Data CorruptionsRui Yang, Jie Wang, Guoping Wu, Bin LiNeurIPS 2024 · 被引用 11 次
- Diverse Projection Ensembles for Distributional Reinforcement LearningMoritz Akiya Zanger, Wendelin Boehmer, Matthijs T. J. SpaanICLR 2024 · 被引用 9 次
- Bayesian Exploration NetworksMattie Fellows, Brandon Kaplowitz, Christian Schröder de Witt, Shimon WhitesonICML 2024 · 被引用 4 次
- Epistemic Bellman OperatorsPascal R. van der Vaart, Matthijs T. J. Spaan, Neil Yorke-SmithAAAI 2025 · 被引用 2 次
它引用的顶会 Paper3
- VariBAD: A Very Good Method for Bayes-Adaptive Deep RL via Meta-LearningLuisa M. Zintgraf, Kyriacos Shiarlis, Maximilian Igl, Sebastian Schulze 等ICLR 2020 · 被引用 315 次
- Conservative Uncertainty Estimation By Fitting Prior NetworksKamil Ciosek, Vincent Fortuin, Ryota Tomioka, Katja Hofmann 等ICLR 2020 · 被引用 65 次
- Geometric Insights into the Convergence of Nonlinear TD LearningDavid Brandfonbrener, Joan BrunaICLR 2020 · 被引用 18 次
相关 Paper
- Principled Exploration via Optimistic Bootstrapping and Backward InductionChenjia Bai, Lingxiao Wang, Lei Han, Jianye Hao 等ICML 2021 · 被引用 46 次
- ReLU to the Rescue: Improve Your On-Policy Actor-Critic with Positive AdvantagesAndrew Jesson, Chris Lu, Gunshi Gupta, Nicolas Beltran-Velez 等ICML 2024 · 被引用 11 次
- Parameterized Projected Bellman OperatorThéo Vincent, Alberto Maria Metelli, Boris Belousov, Jan Peters 等AAAI 2024 · 被引用 6 次
- Sample Efficient Deep Reinforcement Learning via Uncertainty EstimationVincent Mai, Kaustubh Mani, Liam PaullICLR 2022 · 被引用 53 次
- SUNRISE: A Simple Unified Framework for Ensemble Learning in Deep Reinforcement LearningKimin Lee, Michael Laskin, Aravind Srinivas, Pieter AbbeelICML 2021 · 被引用 239 次
