A Distributional Analogue to the Successor Representation
Harley Wiltzer, Jesse Farebrother, Arthur Gretton, Yunhao Tang, André Barreto, Will Dabney, Marc G. Bellemare, Mark Rowland
摘要
This paper contributes a new approach for distributional reinforcement learning which elucidates a clean separation of transition structure and reward in the learning process. Analogous to how the successor representation (SR) describes the expected consequences of behaving according to a given policy, our distributional successor measure (SM) describes the distributional consequences of this behaviour. We formulate the distributional SM as a distribution over distributions and provide theory connecting it with distributional and model-based reinforcement learning. Moreover, we propose an algorithm that learns the distributional SM from data by minimizing a two-level maximum mean discrepancy. Key to our method are a number of algorithmic techniques that are independently valuable for learning generative models of state. As an illustration of the usefulness of the distributional SM, we show that it enables zero-shot risk-sensitive policy evaluation in a way that was not previously possible.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Foundations of Multivariate Distributional Reinforcement LearningHarley Wiltzer, Jesse Farebrother, Arthur Gretton, Mark RowlandNeurIPS 2024 · 被引用 21 次
- Towards Robust Zero-Shot Reinforcement LearningKexin Zheng, Lauriane Teyssier, Yinan Zheng, Yu Luo 等NeurIPS 2025 · 被引用 7 次
- A Finite Sample Analysis of Distributional TD Learning with Linear Function ApproximationYang Peng, Kaicheng Jin, Liangyu Zhang, Zhihua ZhangNeurIPS 2025 · 被引用 6 次
- Reward-Aware Proto-Representations in Reinforcement LearningHon Tik Tse, Siddarth Chandrasekar, Marlos C. MachadoNeurIPS 2025 · 被引用 6 次
- Convergence Theorems for Entropy-Regularized and Distributional Reinforcement LearningYash Jhaveri, Harley Wiltzer, Patrick Shafto, Marc G. Bellemare 等NeurIPS 2025 · 被引用 3 次
它引用的顶会 Paper19
- Contrastive Learning as Goal-Conditioned Reinforcement LearningBenjamin Eysenbach, Tianjun Zhang, Sergey Levine, Ruslan SalakhutdinovNeurIPS 2022 · 被引用 331 次
- Count-Based Exploration with the Successor RepresentationMarlos C. Machado, Marc G. Bellemare, Michael BowlingAAAI 2020 · 被引用 206 次
- Learning One Representation to Optimize All RewardsAhmed Touati, Yann OllivierNeurIPS 2021 · 被引用 140 次
- Einops: Clear and Reliable Tensor Manipulations with Einstein-like NotationAlex RogozhnikovICLR 2022 · 被引用 124 次
- Reinforcement Learning from Passive Data via Latent IntentionsDibya Ghosh, Chethan Anand Bhateja, Sergey LevineICML 2023 · 被引用 69 次
相关 Paper
- Distributional Model Equivalence for Risk-Sensitive Reinforcement LearningTyler Kastner, Murat A. Erdogdu, Amir-massoud FarahmandNeurIPS 2023 · 被引用 9 次
- A Deep Reinforcement Learning Approach to Marginalized Importance Sampling with the Successor RepresentationScott Fujimoto, David Meger, Doina PrecupICML 2021 · 被引用 17 次
- Distributional Successor Features Enable Zero-Shot Policy OptimizationChuning Zhu, Xinqi Wang, Tyler Han, Simon S. Du 等NeurIPS 2024 · 被引用 11 次
- Provable Risk-Sensitive Distributional Reinforcement Learning with General Function ApproximationYu Chen, Xiangcheng Zhang, Siwei Wang, Longbo HuangICML 2024 · 被引用 3 次
- Near-Minimax-Optimal Distributional Reinforcement Learning with a Generative ModelMark Rowland, Kevin Kevin Li, Rémi Munos, Clare Lyle 等NeurIPS 2024 · 被引用 9 次
