Stochastically Dominant Distributional Reinforcement Learning
John D. Martin, Michal Lyskawinski, Xiaohu Li, Brendan J. Englot
Abstract
We describe a new approach for managing aleatoric uncertainty in the Reinforcement Learning (RL) paradigm. Instead of selecting actions according to a single statistic, we propose a distributional method based on the second-order stochastic dominance (SSD) relation. This compares the inherent dispersion of random returns induced by actions, producing a more comprehensive and robust evaluation of the environment's uncertainty. The necessary conditions for SSD require estimators to predict accurate second moments. To accommodate this, we map the distributional RL problem to a Wasserstein gradient flow, treating the distributional Bellman residual as a potential energy functional. We propose a particle-based algorithm for which we prove optimality and convergence. Our experiments characterize the algorithm performance and demonstrate how uncertainty and performance are better balanced using an ssd policy than with other risk measures.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2412cfee-2605-4f7b-8fdc-5a6ea45cf986Cited by top-tier papers4
- Distributional Hamilton-Jacobi-Bellman Equations for Continuous-Time Reinforcement LearningHarley E. Wiltzer, David Meger, Marc G. BellemareICML 2022 · 18 citations
- Bayesian Distributional Policy GradientsLuchen Li, A. Aldo FaisalAAAI 2021 · 11 citations
- Enhancing Value Function Estimation through First-Order State-Action Dynamics in Offline Reinforcement LearningYun-Hsuan Lien, Ping-Chun Hsieh, Tzu-Mao Li, Yu-Shuen WangICML 2024 · 4 citations
- Learning with Stochastic OrdersCarles Domingo-Enrich, Yair Schiff, Youssef MrouehICLR 2023
Builds on1
Related papers
- Pitfall of Optimism: Distributional Reinforcement Learning by Randomizing Risk CriterionTaehyun Cho, Seungyub Han, Heesoo Lee, Kyungjae Lee et al.NeurIPS 2023 · 10 citations
- Wasserstein Actor-Critic: Directed Exploration via Optimism for Continuous-Actions ControlAmarildo Likmeta, Matteo Sacco, Alberto Maria Metelli, Marcello RestelliAAAI 2023 · 7 citations
- Diverse Projection Ensembles for Distributional Reinforcement LearningMoritz Akiya Zanger, Wendelin Boehmer, Matthijs T. J. SpaanICLR 2024 · 9 citations
- Taming Aleatoric Impulse in Off-Policy Reinforcement LearningZhouyang Yu, Guojian Zhan, Yang Guan, Jingliang Duan et al.ICML 2026
- Value FlowsPerry Dong, Chongyi Zheng, Chelsea Finn, Dorsa Sadigh et al.ICLR 2026 · 13 citations
