Understanding Self-Predictive Learning for Reinforcement Learning
Yunhao Tang, Zhaohan Daniel Guo, Pierre Harvey Richemond, Bernardo Ávila Pires, Yash Chandak, Rémi Munos, Mark Rowland, Mohammad Gheshlaghi Azar, Charline Le Lan, Clare Lyle, András György, Shantanu Thakoor
Abstract
We study the learning dynamics of self-predictive learning for reinforcement learning, a family of algorithms that learn representations by minimizing the prediction error of their own future latent representations. Despite its recent empirical success, such algorithms have an apparent defect: trivial representations (such as constants) minimize the prediction error, yet it is obviously undesirable to converge to such solutions. Our central insight is that careful designs of the optimization dynamics are critical to learning meaningful representations. We identify that a faster paced optimization of the predictor and semi-gradient updates on the representation, are crucial to preventing the representation collapse. Then in an idealized setup, we show self-predictive learning dynamics carries out spectral decomposition on the state transition matrix, effectively capturing information of the transition dynamics. Building on the theoretical insights, we propose bidirectional self-predictive learning, a novel self-predictive algorithm that learns two representations simultaneously. We examine the robustness of our theoretical insights with a number of small-scale experiments and showcase the promise of the novel representation learning algorithm with large-scale experiments. We present the first attempt at understanding self-predictive learning for RL, through a theoretical lens. In an idealized setting, we identify key elements to ensure that the selfpredictive algorithm avoids collapse and learns meaningful representations. We make the following theoretical and algorithmic contributions. Key algorithmic elements to prevent collapse. We identify two key algorithmic components: (1) the two time-scale optimization of the transition function P and representation
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext aa73ec94-d46d-4952-9286-93ae7dd05901Cited by top-tier papers25
- Bridging State and History Representations: Understanding Self-Predictive RLTianwei Ni, Benjamin Eysenbach, Erfan Seyedsalehi, Michel Ma et al.ICLR 2024 · 50 citations
- Cross-Domain Policy Adaptation by Capturing Representation MismatchJiafei Lyu, Chenjia Bai, Jingwen Yang, Zongqing Lu et al.ICML 2024 · 30 citations
- Inference via Interpolation: Contrastive Representations Provably Enable Planning and InferenceBenjamin Eysenbach, Vivek Myers, Ruslan Salakhutdinov, Sergey LevineNeurIPS 2024 · 23 citations
- TD-JEPA: Latent-predictive Representations for Zero-Shot Reinforcement LearningMarco Bagatella, Matteo Pirotta, Ahmed Touati, Alessandro Lazaric et al.ICLR 2026 · 19 citations
- Predictive auxiliary objectives in deep RL mimic learning in the brainChing Fang, Kim StachenfeldICLR 2024 · 16 citations
Builds on10
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Data-Efficient Reinforcement Learning with Self-Predictive RepresentationsMax Schwarzer, Ankesh Anand, Rishab Goel, R. Devon Hjelm et al.ICLR 2021 · 399 citations
- Understanding self-supervised learning dynamics without contrastive pairsYuandong Tian, Xinlei Chen, Surya GanguliICML 2021 · 338 citations
- Bootstrap Latent-Predictive Representations for Multitask Reinforcement LearningZhaohan Daniel Guo, Bernardo Ávila Pires, Bilal Piot, Jean-Bastien Grill et al.ICML 2020 · 153 citations
- Learning One Representation to Optimize All RewardsAhmed Touati, Yann OllivierNeurIPS 2021 · 140 citations
Related papers
- Spectral Decomposition Representation for Reinforcement LearningTongzheng Ren, Tianjun Zhang, Lisa Lee, Joseph E. Gonzalez et al.ICLR 2023 · 1 citation
- Towards a better understanding of representation dynamics under TD-learningYunhao Tang, Rémi MunosICML 2023 · 3 citations
- Spectral Bellman Method: Unifying Representation and Exploration in RLOfir Nabati, Bo Dai, Shie Mannor, Guy TennenholtzICLR 2026 · 3 citations
- On the Importance of Feature Decorrelation for Unsupervised Representation Learning in Reinforcement LearningHojoon Lee, Koanho Lee, Dongyoon Hwang, Hyunho Lee et al.ICML 2023 · 11 citations
- Representations and Exploration for Deep Reinforcement Learning using Singular Value DecompositionYash Chandak, Shantanu Thakoor, Zhaohan Daniel Guo, Yunhao Tang et al.ICML 2023 · 6 citations
