Understanding Self-Predictive Learning for Reinforcement Learning
Yunhao Tang, Zhaohan Daniel Guo, Pierre Harvey Richemond, Bernardo Ávila Pires, Yash Chandak, Rémi Munos, Mark Rowland, Mohammad Gheshlaghi Azar, Charline Le Lan, Clare Lyle, András György, Shantanu Thakoor
摘要
We study the learning dynamics of self-predictive learning for reinforcement learning, a family of algorithms that learn representations by minimizing the prediction error of their own future latent representations. Despite its recent empirical success, such algorithms have an apparent defect: trivial representations (such as constants) minimize the prediction error, yet it is obviously undesirable to converge to such solutions. Our central insight is that careful designs of the optimization dynamics are critical to learning meaningful representations. We identify that a faster paced optimization of the predictor and semi-gradient updates on the representation, are crucial to preventing the representation collapse. Then in an idealized setup, we show self-predictive learning dynamics carries out spectral decomposition on the state transition matrix, effectively capturing information of the transition dynamics. Building on the theoretical insights, we propose bidirectional self-predictive learning, a novel self-predictive algorithm that learns two representations simultaneously. We examine the robustness of our theoretical insights with a number of small-scale experiments and showcase the promise of the novel representation learning algorithm with large-scale experiments. We present the first attempt at understanding self-predictive learning for RL, through a theoretical lens. In an idealized setting, we identify key elements to ensure that the selfpredictive algorithm avoids collapse and learns meaningful representations. We make the following theoretical and algorithmic contributions. Key algorithmic elements to prevent collapse. We identify two key algorithmic components: (1) the two time-scale optimization of the transition function P and representation
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper25
- Bridging State and History Representations: Understanding Self-Predictive RLTianwei Ni, Benjamin Eysenbach, Erfan Seyedsalehi, Michel Ma 等ICLR 2024 · 被引用 50 次
- Cross-Domain Policy Adaptation by Capturing Representation MismatchJiafei Lyu, Chenjia Bai, Jingwen Yang, Zongqing Lu 等ICML 2024 · 被引用 30 次
- Inference via Interpolation: Contrastive Representations Provably Enable Planning and InferenceBenjamin Eysenbach, Vivek Myers, Ruslan Salakhutdinov, Sergey LevineNeurIPS 2024 · 被引用 23 次
- TD-JEPA: Latent-predictive Representations for Zero-Shot Reinforcement LearningMarco Bagatella, Matteo Pirotta, Ahmed Touati, Alessandro Lazaric 等ICLR 2026 · 被引用 19 次
- Predictive auxiliary objectives in deep RL mimic learning in the brainChing Fang, Kim StachenfeldICLR 2024 · 被引用 16 次
它引用的顶会 Paper10
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- Data-Efficient Reinforcement Learning with Self-Predictive RepresentationsMax Schwarzer, Ankesh Anand, Rishab Goel, R. Devon Hjelm 等ICLR 2021 · 被引用 399 次
- Understanding self-supervised learning dynamics without contrastive pairsYuandong Tian, Xinlei Chen, Surya GanguliICML 2021 · 被引用 338 次
- Bootstrap Latent-Predictive Representations for Multitask Reinforcement LearningZhaohan Daniel Guo, Bernardo Ávila Pires, Bilal Piot, Jean-Bastien Grill 等ICML 2020 · 被引用 153 次
- Learning One Representation to Optimize All RewardsAhmed Touati, Yann OllivierNeurIPS 2021 · 被引用 140 次
相关 Paper
- Spectral Decomposition Representation for Reinforcement LearningTongzheng Ren, Tianjun Zhang, Lisa Lee, Joseph E. Gonzalez 等ICLR 2023 · 被引用 1 次
- Towards a better understanding of representation dynamics under TD-learningYunhao Tang, Rémi MunosICML 2023 · 被引用 3 次
- Spectral Bellman Method: Unifying Representation and Exploration in RLOfir Nabati, Bo Dai, Shie Mannor, Guy TennenholtzICLR 2026 · 被引用 3 次
- On the Importance of Feature Decorrelation for Unsupervised Representation Learning in Reinforcement LearningHojoon Lee, Koanho Lee, Dongyoon Hwang, Hyunho Lee 等ICML 2023 · 被引用 11 次
- Representations and Exploration for Deep Reinforcement Learning using Singular Value DecompositionYash Chandak, Shantanu Thakoor, Zhaohan Daniel Guo, Yunhao Tang 等ICML 2023 · 被引用 6 次
