Foundations of Multivariate Distributional Reinforcement Learning
Harley Wiltzer, Jesse Farebrother, Arthur Gretton, Mark Rowland
摘要
In reinforcement learning (RL), the consideration of multivariate reward signals has led to fundamental advancements in multi-objective decision-making, transfer learning, and representation learning. This work introduces the first oracle-free and computationally-tractable algorithms for provably convergent multivariate distributional dynamic programming and temporal difference learning. Our convergence rates match the familiar rates in the scalar reward setting, and additionally provide new insights into the fidelity of approximate return distribution representations as a function of the reward dimension. Surprisingly, when the reward dimension is larger than , we show that standard analysis of categorical TD learning fails, which we resolve with a novel projection onto the space of mass- signed measures. Finally, with the aid of our technical results and simulations, we identify tradeoffs between distribution representations that influence the performance of multivariate distributional RL in practice.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- A Finite Sample Analysis of Distributional TD Learning with Linear Function ApproximationYang Peng, Kaicheng Jin, Liangyu Zhang, Zhihua ZhangNeurIPS 2025 · 被引用 6 次
- Convergence Theorems for Entropy-Regularized and Distributional Reinforcement LearningYash Jhaveri, Harley Wiltzer, Patrick Shafto, Marc G. Bellemare 等NeurIPS 2025 · 被引用 3 次
- Distributional Inverse Reinforcement LearningFeiyang Wu, Ye Zhao, Anqi WuICML 2026 · 被引用 1 次
- Intrinsic Benefits of Categorical Distributional Loss: Uncertainty-aware Regularized Exploration in Reinforcement LearningKe Sun, Yingnan Zhao, Enze Shi, Yafei Wang 等NeurIPS 2025 · 被引用 1 次
- Compositional Planning with Jumpy World ModelsJesse Farebrother, Matteo Pirotta, Andrea Tirinzoni, Marc Bellemare 等ICML 2026 · 被引用 1 次
它引用的顶会 Paper9
- DFAC Framework: Factorizing the Value Function via Quantile Mixture for Multi-Agent Distributional Q-LearningWei-Fang Sun, Cheng-Kuang Lee, Chun-Yi LeeICML 2021 · 被引用 56 次
- Distributional Pareto-Optimal Multi-Objective Reinforcement LearningXin-Qiang Cai, Pushi Zhang, Li Zhao, Jiang Bian 等NeurIPS 2023 · 被引用 46 次
- Distributional Reinforcement Learning via Moment MatchingThanh Nguyen-Tang, Sunil Gupta, Svetha VenkateshAAAI 2021 · 被引用 44 次
- Distributional Reinforcement Learning for Multi-Dimensional Reward FunctionsPushi Zhang, Xiaoyu Chen, Li Zhao, Wei Xiong 等NeurIPS 2021 · 被引用 33 次
- Risk-Aware Transfer in Reinforcement Learning using Successor FeaturesMichael Gimelfarb, André Barreto, Scott Sanner, Chi-Guhn LeeNeurIPS 2021 · 被引用 25 次
相关 Paper
- Distributional Bellman Operators over Mean EmbeddingsLi Kevin Wenliang, Grégoire Delétang, Matthew Aitchison, Marcus Hutter 等ICML 2024 · 被引用 5 次
- A Differential Perspective on Distributional Reinforcement LearningJuan Sebastian Rojas, Chi-Guhn LeeAAAI 2026 · 被引用 4 次
- Multivariate Distributional Reinforcement Learning Using Sliced DivergencesBaptiste Debes, Tinne TuytelaarsICML 2026
- Intersectional Fairness in Reinforcement Learning with Large State and Constraint SpacesEric Eaton, Marcel Hussing, Michael Kearns, Aaron Roth 等ICML 2025
- Statistical Efficiency of Distributional Temporal Difference LearningYang Peng, Liangyu Zhang, Zhihua ZhangNeurIPS 2024 · 被引用 8 次
