Stochastically Dominant Distributional Reinforcement Learning
John D. Martin, Michal Lyskawinski, Xiaohu Li, Brendan J. Englot
摘要
We describe a new approach for managing aleatoric uncertainty in the Reinforcement Learning (RL) paradigm. Instead of selecting actions according to a single statistic, we propose a distributional method based on the second-order stochastic dominance (SSD) relation. This compares the inherent dispersion of random returns induced by actions, producing a more comprehensive and robust evaluation of the environment's uncertainty. The necessary conditions for SSD require estimators to predict accurate second moments. To accommodate this, we map the distributional RL problem to a Wasserstein gradient flow, treating the distributional Bellman residual as a potential energy functional. We propose a particle-based algorithm for which we prove optimality and convergence. Our experiments characterize the algorithm performance and demonstrate how uncertainty and performance are better balanced using an ssd policy than with other risk measures.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Distributional Hamilton-Jacobi-Bellman Equations for Continuous-Time Reinforcement LearningHarley E. Wiltzer, David Meger, Marc G. BellemareICML 2022 · 被引用 18 次
- Bayesian Distributional Policy GradientsLuchen Li, A. Aldo FaisalAAAI 2021 · 被引用 11 次
- Enhancing Value Function Estimation through First-Order State-Action Dynamics in Offline Reinforcement LearningYun-Hsuan Lien, Ping-Chun Hsieh, Tzu-Mao Li, Yu-Shuen WangICML 2024 · 被引用 4 次
- Learning with Stochastic OrdersCarles Domingo-Enrich, Yair Schiff, Youssef MrouehICLR 2023
它引用的顶会 Paper1
相关 Paper
- Pitfall of Optimism: Distributional Reinforcement Learning by Randomizing Risk CriterionTaehyun Cho, Seungyub Han, Heesoo Lee, Kyungjae Lee 等NeurIPS 2023 · 被引用 10 次
- Wasserstein Actor-Critic: Directed Exploration via Optimism for Continuous-Actions ControlAmarildo Likmeta, Matteo Sacco, Alberto Maria Metelli, Marcello RestelliAAAI 2023 · 被引用 7 次
- Diverse Projection Ensembles for Distributional Reinforcement LearningMoritz Akiya Zanger, Wendelin Boehmer, Matthijs T. J. SpaanICLR 2024 · 被引用 9 次
- Taming Aleatoric Impulse in Off-Policy Reinforcement LearningZhouyang Yu, Guojian Zhan, Yang Guan, Jingliang Duan 等ICML 2026
- Value FlowsPerry Dong, Chongyi Zheng, Chelsea Finn, Dorsa Sadigh 等ICLR 2026 · 被引用 13 次
