Action Gaps and Advantages in Continuous-Time Distributional Reinforcement Learning
Harley Wiltzer, Marc G. Bellemare, David Meger, Patrick Shafto, Yash Jhaveri
Abstract
When decisions are made at high frequency, traditional reinforcement learning (RL) methods struggle to accurately estimate action values. In turn, their performance is inconsistent and often poor. Whether the performance of distributional RL (DRL) agents suffers similarly, however, is unknown. In this work, we establish that DRL agents are sensitive to the decision frequency. We prove that action-conditioned return distributions collapse to their underlying policy's return distribution as the decision frequency increases. We quantify the rate of collapse of these return distributions and exhibit that their statistics collapse at different rates. Moreover, we define distributional perspectives on action gaps and advantages. In particular, we introduce the superiority as a probabilistic generalization of the advantage -- the core object of approaches to mitigating performance issues in high-frequency value-based RL. In addition, we build a superiority-based DRL algorithm. Through simulations in an option-trading domain, we validate that proper modeling of the superiority distribution produces improved controllers at high decision frequencies.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b3146b2d-bcac-4c0d-8b9d-bfebf96349cfCited by top-tier papers2
- Convergence Theorems for Entropy-Regularized and Distributional Reinforcement LearningYash Jhaveri, Harley Wiltzer, Patrick Shafto, Marc G. Bellemare et al.NeurIPS 2025 · 3 citations
- Intrinsic Benefits of Categorical Distributional Loss: Uncertainty-aware Regularized Exploration in Reinforcement LearningKe Sun, Yingnan Zhao, Enze Shi, Yafei Wang et al.NeurIPS 2025 · 1 citation
Builds on8
- Distributional Reinforcement Learning for Risk-Sensitive PoliciesShiau Hong Lim, Ilyas MalikNeurIPS 2022 · 54 citations
- Policy Optimization for Continuous Reinforcement LearningHanyang Zhao, Wenpin Tang, David D. YaoNeurIPS 2023 · 47 citations
- The Phenomenon of Policy ChurnTom Schaul, André Barreto, John Quan, Georg OstrovskiNeurIPS 2022 · 38 citations
- Direct Advantage EstimationHsiao-Ru Pan, Nico Gürtler, Alexander Neitz, Bernhard SchölkopfNeurIPS 2022 · 20 citations
- Distributional Hamilton-Jacobi-Bellman Equations for Continuous-Time Reinforcement LearningHarley E. Wiltzer, David Meger, Marc G. BellemareICML 2022 · 18 citations
Related papers
- A Differential Perspective on Distributional Reinforcement LearningJuan Sebastian Rojas, Chi-Guhn LeeAAAI 2026 · 4 citations
- Robust Action Gap Increasing with Clipped Advantage LearningZhe Zhang, Yaozhong Gan, Xiaoyang TanAAAI 2022 · 3 citations
- The Statistical Benefits of Quantile Temporal-Difference Learning for Value EstimationMark Rowland, Yunhao Tang, Clare Lyle, Rémi Munos et al.ICML 2023 · 13 citations
- Smoothing Advantage LearningYaozhong Gan, Zhe Zhang, Xiaoyang TanAAAI 2022 · 3 citations
- Distributional Bellman Operators over Mean EmbeddingsLi Kevin Wenliang, Grégoire Delétang, Matthew Aitchison, Marcus Hutter et al.ICML 2024 · 5 citations
