Distributional Reinforcement Learning via Moment Matching
Thanh Nguyen-Tang, Sunil Gupta, Svetha Venkatesh
摘要
We consider the problem of learning a set of probability distributions from the empirical Bellman dynamics in distributional reinforcement learning (RL), a class of state-of-the-art methods that estimate the distribution, as opposed to only the expectation, of the total return. We formulate a method that learns a finite set of statistics from each return distribution via neural networks, as in the distributional RL literature. Existing distributional RL methods however constrain the learned statistics to predefined functional forms of the return distribution which is both restrictive in representation and difficult in maintaining the predefined statistics. Instead, we learn unrestricted statistics, i.e., deterministic (pseudo-)samples, of the return distribution by leveraging a technique from hypothesis testing known as maximum mean discrepancy (MMD), which leads to a simpler objective amenable to backpropagation. Our method can be interpreted as implicitly matching all orders of moments between a return distribution and its Bellman target. We establish sufficient conditions for the contraction of the distributional Bellman operator and provide finite-sample analysis for the deterministic samples in distribution approximation. Experiments on the suite of Atari games show that our method outperforms the standard distributional RL baselines and sets a new record in the Atari games for non-distributed agents.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- Conservative Offline Distributional Reinforcement LearningYecheng Jason Ma, Dinesh Jayaraman, Osbert BastaniNeurIPS 2021 · 被引用 118 次
- Distributional Reinforcement Learning for Multi-Dimensional Reward FunctionsPushi Zhang, Xiaoyu Chen, Li Zhao, Wei Xiong 等NeurIPS 2021 · 被引用 33 次
- Foundations of Multivariate Distributional Reinforcement LearningHarley Wiltzer, Jesse Farebrother, Arthur Gretton, Mark RowlandNeurIPS 2024 · 被引用 21 次
- The Statistical Benefits of Quantile Temporal-Difference Learning for Value EstimationMark Rowland, Yunhao Tang, Clare Lyle, Rémi Munos 等ICML 2023 · 被引用 13 次
- A Distributional Analogue to the Successor RepresentationHarley Wiltzer, Jesse Farebrother, Arthur Gretton, Yunhao Tang 等ICML 2024 · 被引用 11 次
相关 Paper
- Distributional Reinforcement Learning with Regularized Wasserstein LossKe Sun, Yingnan Zhao, Wulong Liu, Bei Jiang 等NeurIPS 2024 · 被引用 2 次
- Multivariate Distributional Reinforcement Learning Using Sliced DivergencesBaptiste Debes, Tinne TuytelaarsICML 2026
- Diverse Projection Ensembles for Distributional Reinforcement LearningMoritz Akiya Zanger, Wendelin Boehmer, Matthijs T. J. SpaanICLR 2024 · 被引用 9 次
- Statistical Efficiency of Distributional Temporal Difference LearningYang Peng, Liangyu Zhang, Zhihua ZhangNeurIPS 2024 · 被引用 8 次
- Distributional Bellman Operators over Mean EmbeddingsLi Kevin Wenliang, Grégoire Delétang, Matthew Aitchison, Marcus Hutter 等ICML 2024 · 被引用 5 次
