Bayesian Distributional Policy Gradients
Luchen Li, A. Aldo Faisal
Abstract
Distributional Reinforcement Learning (RL) maintains the entire probability distribution of the reward-to-go, i.e. the return, providing more learning signals that account for the uncertainty associated with policy performance, which may be beneficial for trading off exploration and exploitation and policy learning in general. Previous works in distributional RL focused mainly on computing the state-action-return distributions, here we model the state-return distributions. This enables us to translate successful conventional RL algorithms that are based on state values into distributional RL. We formulate the distributional Bellman operation as an inference-based auto-encoding process that minimises Wasserstein metrics between target/model return distributions. The proposed algorithm, BDPG (Bayesian Distributional Policy Gradients), uses adversarial training in joint-contrastive learning to estimate a variational posterior from the returns. Moreover, we can now interpret the return prediction uncertainty as an information gain, which allows to obtain a new curiosity measure that helps BDPG steer exploration actively and efficiently. We demonstrate in a suite of Atari 2600 games and MuJoCo tasks, including well known hard-exploration challenges, how BDPG learns generally faster and with higher asymptotic performance than reference distributional RL algorithms.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7632ab98-83da-4032-ae57-1d49daf7b896Cited by top-tier papers2
- Distributional Offline Policy Evaluation with Predictive Error GuaranteesRunzhe Wu, Masatoshi Uehara, Wen SunICML 2023 · 19 citations
- Beyond Noisy-TVs: Noise-Robust Exploration Via Learning Progress MonitoringZhibo Hou, Zhiyu An, Wan DuICLR 2026 · 3 citations
Builds on4
- Planning to Explore via Self-Supervised World ModelsRamanan Sekar, Oleh Rybkin, Kostas Daniilidis, Pieter Abbeel et al.ICML 2020 · 489 citations
- Controlling Overestimation Bias with Truncated Mixture of Continuous Distributional Quantile CriticsArsenii Kuznetsov, Pavel Shvechikov, Alexander Grishin, Dmitry P. VetrovICML 2020 · 266 citations
- Ready Policy One: World Building Through Active LearningPhilip J. Ball, Jack Parker-Holder, Aldo Pacchiano, Krzysztof Choromanski et al.ICML 2020 · 52 citations
- Stochastically Dominant Distributional Reinforcement LearningJohn D. Martin, Michal Lyskawinski, Xiaohu Li, Brendan J. EnglotICML 2020 · 26 citations
Related papers
- Distributional value gradients for stochastic environmentsBaptiste Debes, Tinne TuytelaarsICLR 2026
- Wasserstein Actor-Critic: Directed Exploration via Optimism for Continuous-Actions ControlAmarildo Likmeta, Matteo Sacco, Alberto Maria Metelli, Marcello RestelliAAAI 2023 · 7 citations
- Distributional Meta-Gradient Reinforcement LearningHaiyan Yin, Shuicheng Yan, Zhongwen XuICLR 2023
- Distributional Reinforcement Learning with Regularized Wasserstein LossKe Sun, Yingnan Zhao, Wulong Liu, Bei Jiang et al.NeurIPS 2024 · 2 citations
- Distributional Reinforcement Learning with Monotonic SplinesYudong Luo, Guiliang Liu, Haonan Duan, Oliver Schulte et al.ICLR 2022 · 18 citations
