Bayesian Exploration Networks
Mattie Fellows, Brandon Kaplowitz, Christian Schröder de Witt, Shimon Whiteson
Abstract
Bayesian reinforcement learning (RL) offers a principled and elegant approach for sequential decision making under uncertainty. Most notably, Bayesian agents do not face an exploration/exploitation dilemma, a major pathology of frequentist methods. However theoretical understanding of model-free approaches is lacking. In this paper, we introduce a novel Bayesian model-free formulation and the first analysis showing that model-free approaches can yield Bayes-optimal policies. We show all existing model-free approaches make approximations that yield policies that can be arbitrarily Bayes-suboptimal. As a first step towards model-free Bayes optimality, we introduce the Bayesian exploration network (BEN) which uses normalising flows to model both the aleatoric uncertainty (via density estimation) and epistemic uncertainty (via variational inference) in the Bellman operator. In the limit of complete optimisation, BEN learns true Bayes-optimal policies, but like in variational expectation-maximisation, partial optimisation renders our approach tractable. Empirical results demonstrate that BEN can learn true Bayes-optimal policies in tasks where existing model-free approaches fail.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bcc43df2-ed55-4b46-bea3-407acaa6fc9eCited by top-tier papers1
Ask how each one uses itBuilds on7
- VariBAD: A Very Good Method for Bayes-Adaptive Deep RL via Meta-LearningLuisa M. Zintgraf, Kyriacos Shiarlis, Maximilian Igl, Sebastian Schulze et al.ICLR 2020 · 315 citations
- SurVAE Flows: Surjections to Bridge the Gap between VAEs and FlowsDidrik Nielsen, Priyank Jaini, Emiel Hoogeboom, Ole Winther et al.NeurIPS 2020 · 100 citations
- Improving Generalization in Meta-learning via Task AugmentationHuaxiu Yao, Long-Kai Huang, Linjun Zhang, Ying Wei et al.ICML 2021 · 98 citations
- Exploration in Approximate Hyper-State Space for Meta Reinforcement LearningLuisa M. Zintgraf, Leo Feng, Cong Lu, Maximilian Igl et al.ICML 2021 · 45 citations
- Bayesian Bellman OperatorsMattie Fellows, Kristian Hartikainen, Shimon WhitesonNeurIPS 2021 · 20 citations
Related papers
- Flow-based Recurrent Belief State Learning for POMDPsXiaoyu Chen, Yao Mark Mu, Ping Luo, Shengbo Li et al.ICML 2022 · 26 citations
- Maximum Entropy Reinforcement Learning via Energy-Based Normalizing FlowChen-Hao Chao, Chien Feng, Wei-Fang Sun, Cheng-Kuang Lee et al.NeurIPS 2024 · 29 citations
- Posterior Network: Uncertainty Estimation without OOD Samples via Density-Based Pseudo-CountsBertrand Charpentier, Daniel Zügner, Stephan GünnemannNeurIPS 2020 · 263 citations
- Cooperative Variance Estimation and Bayesian Neural Networks for Disentangling Aleatoric and Epistemic UncertaintiesJiaxiang Yi, Miguel BessaICML 2026 · 2 citations
- Probabilistic Inference in Reinforcement Learning Done RightJean Tarbouriech, Tor Lattimore, Brendan O'DonoghueNeurIPS 2023 · 15 citations
