Prediction and Control in Continual Reinforcement Learning
Nishanth Anand, Doina Precup
摘要
Temporal difference (TD) learning is often used to update the estimate of the value function which is used by RL agents to extract useful policies. In this paper, we focus on value function estimation in continual reinforcement learning. We propose to decompose the value function into two components which update at different timescales: a permanent value function, which holds general knowledge that persists over time, and a transient value function, which allows quick adaptation to new situations. We establish theoretical results showing that our approach is well suited for continual learning and draw connections to the complementary learning systems (CLS) theory from neuroscience. Empirically, this approach improves performance significantly on both prediction and control problems.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Slow and Steady Wins the Race: Maintaining Plasticity with Hare and Tortoise NetworksHojoon Lee, Hyeonseo Cho, Hyunseung Kim, Donghu Kim 等ICML 2024 · 被引用 36 次
- Learning Successor Features the Simple WayRaymond Chua, Arna Ghosh, Christos Kaplanis, Blake A. Richards 等NeurIPS 2024 · 被引用 14 次
- Continual Knowledge Adaptation for Reinforcement LearningJinwu Hu, Zihao Lian, Zhiquan Wen, Chenghao Li 等NeurIPS 2025 · 被引用 8 次
- Tackling Continual Offline RL through Selective Weights Activation on Aligned SpacesJifeng Hu, Sili Huang, Li Shen, Zhejian Yang 等NeurIPS 2025 · 被引用 2 次
- Principled Fast and Meta Knowledge Learners for Continual Reinforcement LearningKe Sun, Hongming Zhang, Jun Jin, Chao Gao 等ICLR 2026 · 被引用 1 次
它引用的顶会 Paper8
- The Primacy Bias in Deep Reinforcement LearningEvgenii Nikishin, Max Schwarzer, Pierluca D'Oro, Pierre-Luc Bacon 等ICML 2022 · 被引用 269 次
- Learning Fast, Learning Slow: A General Continual Learning Method based on Complementary Learning SystemElahe Arani, Fahad Sarfraz, Bahram ZonoozICLR 2022 · 被引用 168 次
- A Definition of Continual Reinforcement LearningDavid Abel, André Barreto, Benjamin Van Roy, Doina Precup 等NeurIPS 2023 · 被引用 167 次
- Understanding and Preventing Capacity Loss in Reinforcement LearningClare Lyle, Mark Rowland, Will DabneyICLR 2022 · 被引用 151 次
- Bootstrapped Meta-LearningSebastian Flennerhag, Yannick Schroecker, Tom Zahavy, Hado van Hasselt 等ICLR 2022 · 被引用 62 次
相关 Paper
- Value Function Decomposition in Markov Recommendation ProcessXiaobei Wang, Shuchang Liu, Qingpeng Cai, Xiang Li 等WWW 2025 · 被引用 4 次
- Modular Lifelong Reinforcement Learning via Neural CompositionJorge A. Mendez, Harm van Seijen, Eric EatonICLR 2022 · 被引用 51 次
- Gamma-Nets: Generalizing Value Estimation over TimescaleCraig Sherstan, Shibhansh Dohare, James MacGlashan, Johannes Günther 等AAAI 2020 · 被引用 14 次
- Foresee then Evaluate: Decomposing Value Estimation with Latent Future PredictionHongyao Tang, Zhaopeng Meng, Guangyong Chen, Pengfei Chen 等AAAI 2021 · 被引用 5 次
- Temporal-Difference Variational Continual LearningLuckeciano Carvalho Melo, Alessandro Abate, Yarin GalNeurIPS 2025 · 被引用 1 次
