VariBAD: A Very Good Method for Bayes-Adaptive Deep RL via Meta-Learning
Luisa M. Zintgraf, Kyriacos Shiarlis, Maximilian Igl, Sebastian Schulze, Yarin Gal, Katja Hofmann, Shimon Whiteson
Abstract
Trading off exploration and exploitation in an unknown environment is key to maximising expected return during learning. A Bayes-optimal policy, which does so optimally, conditions its actions not only on the environment state but on the agent's uncertainty about the environment. Computing a Bayes-optimal policy is however intractable for all but the smallest tasks. In this paper, we introduce variational Bayes-Adaptive Deep RL (variBAD), a way to meta-learn to perform approximate inference in an unknown environment, and incorporate task uncertainty directly during action selection. In a grid-world domain, we illustrate how variBAD performs structured online exploration as a function of task uncertainty. We further evaluate variBAD on MuJoCo domains widely used in meta-RL and show that it achieves higher online return than existing methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 24a56916-20bb-4ea3-b2b2-27db9e9e06d8Cited by top-tier papers93
- Why Generalization in RL is Difficult: Epistemic POMDPs and Implicit Partial ObservabilityDibya Ghosh, Jad Rahme, Aviral Kumar, Amy Zhang et al.NeurIPS 2021 · 176 citations
- Supervised Pretraining Can Learn In-Context Reinforcement LearningJonathan Lee, Annie Xie, Aldo Pacchiano, Yash Chandak et al.NeurIPS 2023 · 170 citations
- Diffusion Model is an Effective Planner and Data Synthesizer for Multi-Task Reinforcement LearningHaoran He, Chenjia Bai, Kang Xu, Zhuoran Yang et al.NeurIPS 2023 · 165 citations
- Recurrent Model-Free RL Can Be a Strong Baseline for Many POMDPsTianwei Ni, Benjamin Eysenbach, Ruslan SalakhutdinovICML 2022 · 162 citations
- Parrot: Data-Driven Behavioral Priors for Reinforcement LearningAvi Singh, Huihan Liu, Gaoyue Zhou, Albert Yu et al.ICLR 2021 · 161 citations
Builds on1
Related papers
- Offline Meta Reinforcement Learning - Identifiability Challenges and Effective Data Collection StrategiesRon Dorfman, Idan Shenfeld, Aviv TamarNeurIPS 2021 · 76 citations
- MetaCARD: Meta-Reinforcement Learning with Task Uncertainty Feedback via Decoupled Context-Aware Reward and Dynamics ComponentsMin Wang, Xin Li, Leiji Zhang, Mingzhong WangAAAI 2024 · 6 citations
- Scalable Bayesian Inverse Reinforcement LearningAlex James Chan, Mihaela van der SchaarICLR 2021 · 11 citations
- ContraBAR: Contrastive Bayes-Adaptive Deep RLEra Choshen, Aviv TamarICML 2023 · 10 citations
- Probabilistic Inference in Reinforcement Learning Done RightJean Tarbouriech, Tor Lattimore, Brendan O'DonoghueNeurIPS 2023 · 15 citations
