Bayesian Exploration Networks
Mattie Fellows, Brandon Kaplowitz, Christian Schröder de Witt, Shimon Whiteson
摘要
Bayesian reinforcement learning (RL) offers a principled and elegant approach for sequential decision making under uncertainty. Most notably, Bayesian agents do not face an exploration/exploitation dilemma, a major pathology of frequentist methods. However theoretical understanding of model-free approaches is lacking. In this paper, we introduce a novel Bayesian model-free formulation and the first analysis showing that model-free approaches can yield Bayes-optimal policies. We show all existing model-free approaches make approximations that yield policies that can be arbitrarily Bayes-suboptimal. As a first step towards model-free Bayes optimality, we introduce the Bayesian exploration network (BEN) which uses normalising flows to model both the aleatoric uncertainty (via density estimation) and epistemic uncertainty (via variational inference) in the Bellman operator. In the limit of complete optimisation, BEN learns true Bayes-optimal policies, but like in variational expectation-maximisation, partial optimisation renders our approach tractable. Empirical results demonstrate that BEN can learn true Bayes-optimal policies in tasks where existing model-free approaches fail.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper7
- VariBAD: A Very Good Method for Bayes-Adaptive Deep RL via Meta-LearningLuisa M. Zintgraf, Kyriacos Shiarlis, Maximilian Igl, Sebastian Schulze 等ICLR 2020 · 被引用 315 次
- SurVAE Flows: Surjections to Bridge the Gap between VAEs and FlowsDidrik Nielsen, Priyank Jaini, Emiel Hoogeboom, Ole Winther 等NeurIPS 2020 · 被引用 100 次
- Improving Generalization in Meta-learning via Task AugmentationHuaxiu Yao, Long-Kai Huang, Linjun Zhang, Ying Wei 等ICML 2021 · 被引用 98 次
- Exploration in Approximate Hyper-State Space for Meta Reinforcement LearningLuisa M. Zintgraf, Leo Feng, Cong Lu, Maximilian Igl 等ICML 2021 · 被引用 45 次
- Bayesian Bellman OperatorsMattie Fellows, Kristian Hartikainen, Shimon WhitesonNeurIPS 2021 · 被引用 20 次
相关 Paper
- Flow-based Recurrent Belief State Learning for POMDPsXiaoyu Chen, Yao Mark Mu, Ping Luo, Shengbo Li 等ICML 2022 · 被引用 26 次
- Maximum Entropy Reinforcement Learning via Energy-Based Normalizing FlowChen-Hao Chao, Chien Feng, Wei-Fang Sun, Cheng-Kuang Lee 等NeurIPS 2024 · 被引用 29 次
- Posterior Network: Uncertainty Estimation without OOD Samples via Density-Based Pseudo-CountsBertrand Charpentier, Daniel Zügner, Stephan GünnemannNeurIPS 2020 · 被引用 263 次
- Cooperative Variance Estimation and Bayesian Neural Networks for Disentangling Aleatoric and Epistemic UncertaintiesJiaxiang Yi, Miguel BessaICML 2026 · 被引用 2 次
- Probabilistic Inference in Reinforcement Learning Done RightJean Tarbouriech, Tor Lattimore, Brendan O'DonoghueNeurIPS 2023 · 被引用 15 次
