Diverse Projection Ensembles for Distributional Reinforcement Learning
Moritz Akiya Zanger, Wendelin Boehmer, Matthijs T. J. Spaan
Abstract
In contrast to classical reinforcement learning (RL), distributional RL algorithms aim to learn the distribution of returns rather than their expected value. Since the nature of the return distribution is generally unknown a priori or arbitrarily complex, a common approach finds approximations within a set of representable, parametric distributions. Typically, this involves a projection of the unconstrained distribution onto the set of simplified distributions. We argue that this projection step entails a strong inductive bias when coupled with neural networks and gradient descent, thereby profoundly impacting the generalization behavior of learned models. In order to facilitate reliable uncertainty estimation through diversity, we study the combination of several different projections and representations in a distributional ensemble. We establish theoretical properties of such projection ensembles and derive an algorithm that uses ensemble disagreement, measured by the average 1-Wasserstein distance, as a bonus for deep exploration. We evaluate our algorithm on the behavior suite benchmark and VizDoom and find that diverse projection ensembles lead to significant performance improvements over existing methods on a variety of tasks with the most pronounced gains in directed exploration problems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 878f4e54-f552-41ae-9778-bbf016be5064Cited by top-tier papers4
- Contextual Similarity Distillation: Ensemble Uncertainties with a Single ModelMoritz Akiya Zanger, Pascal R. Van der Vaart, Wendelin Boehmer, Matthijs T. J. SpaanICLR 2026 · 6 citations
- Compositional Imprecise Probability: A Solution from Graded Monads and Markov CategoriesJack Liell-Cock, Sam StatonPOPL 2025 · 3 citations
- Universal Value-Function UncertaintiesMoritz Akiya Zanger, Max Weltevrede, Yaniv Oren, Pascal R. van der Vaart et al.ICLR 2026 · 1 citation
- Off-Policy Safe Reinforcement Learning with Cost-Constrained Optimistic ExplorationGuopeng Li, Matthijs T. J. Spaan, Julian F. P. KooijICLR 2026
Builds on4
- Behaviour Suite for Reinforcement LearningIan Osband, Yotam Doron, Matteo Hessel, John Aslanides et al.ICLR 2020 · 204 citations
- On the Importance of Exploration for Generalization in Reinforcement LearningYiding Jiang, J. Zico Kolter, Roberta RaileanuNeurIPS 2023 · 48 citations
- Distributional Reinforcement Learning via Moment MatchingThanh Nguyen-Tang, Sunil Gupta, Svetha VenkateshAAAI 2021 · 44 citations
- Bayesian Bellman OperatorsMattie Fellows, Kristian Hartikainen, Shimon WhitesonNeurIPS 2021 · 20 citations
Related papers
- Multivariate Distributional Reinforcement Learning Using Sliced DivergencesBaptiste Debes, Tinne TuytelaarsICML 2026
- Distributional Reinforcement Learning with Regularized Wasserstein LossKe Sun, Yingnan Zhao, Wulong Liu, Bei Jiang et al.NeurIPS 2024 · 2 citations
- Bayesian Distributional Policy GradientsLuchen Li, A. Aldo FaisalAAAI 2021 · 11 citations
- Maximizing Ensemble Diversity in Deep Reinforcement LearningHassam Sheikh, Mariano Phielipp, Ladislau BölöniICLR 2022 · 10 citations
- Ensemble Bootstrapping for Q-LearningOren Peer, Chen Tessler, Nadav Merlis, Ron MeirICML 2021 · 56 citations
