Towards a better understanding of representation dynamics under TD-learning
Yunhao Tang, Rémi Munos
摘要
TD-learning is a foundation reinforcement learning (RL) algorithm for value prediction. Critical to the accuracy of value predictions is the quality of state representations. In this work, we consider the question: how does end-to-end TD-learning impact the representation over time? Complementary to prior work, we provide a set of analysis that sheds further light on the representation dynamics under TD-learning. We first show that when the environments are reversible, end-to-end TD-learning strictly decreases the value approximation error over time. Under further assumptions on the environments, we can connect the representation dynamics with spectral decomposition over the transition matrix. This latter finding establishes fitting multiple value functions from randomly generated rewards as a useful auxiliary task for representation learning, as we empirically validate on both tabular and Atari game suites.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville 等NeurIPS 2021 · 被引用 1,067 次
- Data-Efficient Reinforcement Learning with Self-Predictive RepresentationsMax Schwarzer, Ankesh Anand, Rishab Goel, R. Devon Hjelm 等ICLR 2021 · 被引用 399 次
- Bootstrap Latent-Predictive Representations for Multitask Reinforcement LearningZhaohan Daniel Guo, Bernardo Ávila Pires, Bilal Piot, Jean-Bastien Grill 等ICML 2020 · 被引用 153 次
- The Value-Improvement Path: Towards Better Representations for Reinforcement LearningWill Dabney, André Barreto, Mark Rowland, Robert Dadashi 等AAAI 2021 · 被引用 76 次
- Understanding Self-Predictive Learning for Reinforcement LearningYunhao Tang, Zhaohan Daniel Guo, Pierre Harvey Richemond, Bernardo Ávila Pires 等ICML 2023 · 被引用 46 次
相关 Paper
- Representations for Stable Off-Policy Reinforcement LearningDibya Ghosh, Marc G. BellemareICML 2020 · 被引用 46 次
- Spectral Decomposition Representation for Reinforcement LearningTongzheng Ren, Tianjun Zhang, Lisa Lee, Joseph E. Gonzalez 等ICLR 2023 · 被引用 1 次
- Learning Dynamics and Generalization in Deep Reinforcement LearningClare Lyle, Mark Rowland, Will Dabney, Marta Kwiatkowska 等ICML 2022 · 被引用 40 次
- Value-driven Hindsight ModellingArthur Guez, Fabio Viola, Theophane Weber, Lars Buesing 等NeurIPS 2020 · 被引用 12 次
- Impact of Connectivity on Laplacian Representations in Reinforcement LearningTommaso Giorgi, Pierriccardo Olivieri, Keyue Jiang, Laura Toni 等ICML 2026 · 被引用 1 次
