Prediction and Control in Continual Reinforcement Learning
Nishanth Anand, Doina Precup
Abstract
Temporal difference (TD) learning is often used to update the estimate of the value function which is used by RL agents to extract useful policies. In this paper, we focus on value function estimation in continual reinforcement learning. We propose to decompose the value function into two components which update at different timescales: a permanent value function, which holds general knowledge that persists over time, and a transient value function, which allows quick adaptation to new situations. We establish theoretical results showing that our approach is well suited for continual learning and draw connections to the complementary learning systems (CLS) theory from neuroscience. Empirically, this approach improves performance significantly on both prediction and control problems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0e377d36-df5b-4d17-a424-fa491be5cab5Cited by top-tier papers8
- Slow and Steady Wins the Race: Maintaining Plasticity with Hare and Tortoise NetworksHojoon Lee, Hyeonseo Cho, Hyunseung Kim, Donghu Kim et al.ICML 2024 · 36 citations
- Learning Successor Features the Simple WayRaymond Chua, Arna Ghosh, Christos Kaplanis, Blake A. Richards et al.NeurIPS 2024 · 14 citations
- Continual Knowledge Adaptation for Reinforcement LearningJinwu Hu, Zihao Lian, Zhiquan Wen, Chenghao Li et al.NeurIPS 2025 · 8 citations
- Tackling Continual Offline RL through Selective Weights Activation on Aligned SpacesJifeng Hu, Sili Huang, Li Shen, Zhejian Yang et al.NeurIPS 2025 · 2 citations
- Principled Fast and Meta Knowledge Learners for Continual Reinforcement LearningKe Sun, Hongming Zhang, Jun Jin, Chao Gao et al.ICLR 2026 · 1 citation
Builds on8
- The Primacy Bias in Deep Reinforcement LearningEvgenii Nikishin, Max Schwarzer, Pierluca D'Oro, Pierre-Luc Bacon et al.ICML 2022 · 269 citations
- Learning Fast, Learning Slow: A General Continual Learning Method based on Complementary Learning SystemElahe Arani, Fahad Sarfraz, Bahram ZonoozICLR 2022 · 168 citations
- A Definition of Continual Reinforcement LearningDavid Abel, André Barreto, Benjamin Van Roy, Doina Precup et al.NeurIPS 2023 · 167 citations
- Understanding and Preventing Capacity Loss in Reinforcement LearningClare Lyle, Mark Rowland, Will DabneyICLR 2022 · 151 citations
- Bootstrapped Meta-LearningSebastian Flennerhag, Yannick Schroecker, Tom Zahavy, Hado van Hasselt et al.ICLR 2022 · 62 citations
Related papers
- Value Function Decomposition in Markov Recommendation ProcessXiaobei Wang, Shuchang Liu, Qingpeng Cai, Xiang Li et al.WWW 2025 · 4 citations
- Modular Lifelong Reinforcement Learning via Neural CompositionJorge A. Mendez, Harm van Seijen, Eric EatonICLR 2022 · 51 citations
- Gamma-Nets: Generalizing Value Estimation over TimescaleCraig Sherstan, Shibhansh Dohare, James MacGlashan, Johannes Günther et al.AAAI 2020 · 14 citations
- Foresee then Evaluate: Decomposing Value Estimation with Latent Future PredictionHongyao Tang, Zhaopeng Meng, Guangyong Chen, Pengfei Chen et al.AAAI 2021 · 5 citations
- Temporal-Difference Variational Continual LearningLuckeciano Carvalho Melo, Alessandro Abate, Yarin GalNeurIPS 2025 · 1 citation
