Continual Reinforcement Learning by Planning with Online World Models
Zichen Liu, Guoji Fu, Chao Du, Wee Sun Lee, Min Lin
摘要
Continual reinforcement learning (CRL) refers to a naturalistic setting where an agent needs to endlessly evolve, by trial and error, to solve multiple tasks that are presented sequentially. One of the largest obstacles to CRL is that the agent may forget how to solve previous tasks when learning a new task, known as catastrophic forgetting. In this paper, we propose to address this challenge by planning with online world models. Specifically, we learn a Follow-The-Leader shallow model online to capture the world dynamics, in which we plan using model predictive control to solve a set of tasks specified by any reward functions. The online world model is immune to forgetting by construction with a proven regret bound of O( K 2 D log(T )) under mild assumptions. The planner searches actions solely based on the latest online model, thus forming a FTL Online Agent (OA) that updates incrementally. To assess OA, we further design Continual Bench, a dedicated environment for CRL, and compare with several strong baselines under the same model-planning algorithmic framework. The empirical results show that OA learns continuously to solve new tasks while not forgetting old skills, outperforming agents built on deep world models with various continual learning techniques.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- MEAL: A Benchmark for Continual Multi-Agent Reinforcement LearningTristan Tomilin, Luka van den Boogaard, Samuel Garcin, Constantin Ruhdorfer 等ICML 2026 · 被引用 9 次
- Principled Fast and Meta Knowledge Learners for Continual Reinforcement LearningKe Sun, Hongming Zhang, Jun Jin, Chao Gao 等ICLR 2026 · 被引用 1 次
- Learning Task-Sufficient World Models by Synergizing Agentic Exploration and Structured ModelingFan Feng, Yujia Zheng, Minghao Fu, Yongqiang Chen 等ICML 2026
- Motion Dynamics Learning for Few-Shot Embodied AdaptationSibo He, Weiying Xie, Daixun Li, Junhao Zhong 等ICML 2026
- World Models in Pieces: Structural Certification for General AgentsYikai Lu, Yifei Wu, Xinyu Lu, Tongxin LiICML 2026
它引用的顶会 Paper15
- Leveraging Procedural Generation to Benchmark Reinforcement LearningKarl Cobbe, Christopher Hesse, Jacob Hilton, John SchulmanICML 2020 · 被引用 685 次
- The NetHack Learning EnvironmentHeinrich Küttler, Nantas Nardelli, Alexander H. Miller, Roberta Raileanu 等NeurIPS 2020 · 被引用 251 次
- A Definition of Continual Reinforcement LearningDavid Abel, André Barreto, Benjamin Van Roy, Doina Precup 等NeurIPS 2023 · 被引用 167 次
- Exploring Model-based Planning with Policy NetworksTingwu Wang, Jimmy BaICLR 2020 · 被引用 164 次
- Continual World: A Robotic Benchmark For Continual Reinforcement LearningMaciej Wolczyk, Michal Zajac, Razvan Pascanu, Lukasz Kucinski 等NeurIPS 2021 · 被引用 152 次
相关 Paper
- Continual Predictive Learning from VideosGeng Chen, Wendong Zhang, Han Lu, Siyu Gao 等CVPR 2022 · 被引用 5 次
- Locality Sensitive Sparse Encoding for Learning World Models OnlineZichen Liu, Chao Du, Wee Sun Lee, Min LinICLR 2024 · 被引用 18 次
- Online Fast Adaptation and Knowledge Accumulation (OSAKA): a New Approach to Continual LearningMassimo Caccia, Pau Rodríguez, Oleksiy Ostapenko, Fabrice Normandin 等NeurIPS 2020 · 被引用 83 次
- Same State, Different Task: Continual Reinforcement Learning without InterferenceSamuel Kessler, Jack Parker-Holder, Philip J. Ball, Stefan Zohren 等AAAI 2022 · 被引用 57 次
- The Ideal Continual Learner: An Agent That Never ForgetsLiangzu Peng, Paris Giampouras, René VidalICML 2023 · 被引用 39 次
