A Forget-and-Grow Strategy for Deep Reinforcement Learning Scaling in Continuous Control
Zilin Kang, Chenyuan Hu, Yu Luo, Zhecheng Yuan, Ruijie Zheng, Huazhe Xu
摘要
Deep reinforcement learning for continuous control has recently achieved impressive progress. However, existing methods often suffer from primacy bias-a tendency to overfit early experiences stored in the replay buffer-which limits an RL agent's sample efficiency and generalizability. In contrast, humans are less susceptible to such bias, partly due to infantile amnesia, where the formation of new neurons disrupts early memory traces, leading to the forgetting of initial experiences (Akers et al., 2014) . Inspired by this dual processes of forgetting and growing in neuroscience, in this paper, we propose Forget and Grow (FoG), a new deep RL algorithm with two mechanisms introduced. First, Experience Replay Decay (ER Decay)-"forgetting early experience"-which balances memory by gradually reducing the influence of early experiences. Second, Network Expansion-"growing neural capacity"-which enhances agents' capability to exploit the patterns of existing data by dynamically adding new parameters during training. Empirical results on four major continuous control benchmarks with more than 40 tasks demonstrate the superior performance of FoG against SoTA existing deep RL algorithms, including BRO, SimBa and TD-MPC2.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- XQC: Well-conditioned Optimization Accelerates Deep Reinforcement LearningDaniel Palenicek, Florian Vogt, Joe Watson, Ingmar Posner 等ICLR 2026 · 被引用 20 次
- Entropy Regularizing Activation: Boosting Continuous Control, Large Language Models, and Image Classification with Activation as Entropy ConstraintsZilin Kang, Chonghua Liao, Tingqiang Xu, Huazhe XuICLR 2026 · 被引用 5 次
- The Rank and Gradient Lost in Non-stationarity: Sample Weight Decay for Mitigating Plasticity Loss in Reinforcement LearningZihao Wu, Hongyao Tang, Yi Ma, Jiashun Liu 等ICLR 2026 · 被引用 2 次
- Debiased Model-based Representations for Sample-efficient Continuous ControlJiafei Lyu, Zichuan Lin, Scott Fujimoto, Kai Yang 等ICML 2026
它引用的顶会 Paper20
- TD-MPC2: Scalable, Robust World Models for Continuous ControlNicklas Hansen, Hao Su, Xiaolong WangICLR 2024 · 被引用 388 次
- Revisiting Fundamentals of Experience ReplayWilliam Fedus, Prajit Ramachandran, Rishabh Agarwal, Yoshua Bengio 等ICML 2020 · 被引用 303 次
- Multi-Game Decision TransformersKuang-Huei Lee, Ofir Nachum, Mengjiao Yang, Lisa Lee 等NeurIPS 2022 · 被引用 279 次
- The Primacy Bias in Deep Reinforcement LearningEvgenii Nikishin, Max Schwarzer, Pierluca D'Oro, Pierre-Luc Bacon 等ICML 2022 · 被引用 269 次
- Maxmin Q-learning: Controlling the Estimation Bias of Q-learningQingfeng Lan, Yangchen Pan, Alona Fyshe, Martha WhiteICLR 2020 · 被引用 213 次
相关 Paper
- Neuroplastic Expansion in Deep Reinforcement LearningJiashun Liu, Johan S. Obando-Ceron, Aaron C. Courville, Ling PanICLR 2025
- Sample-Efficient and Safe Deep Reinforcement Learning via Reset Deep Ensemble AgentsWoojun Kim, Yongjae Shin, Jongeui Park, Youngchul SungNeurIPS 2023 · 被引用 20 次
- Self-Composing Policies for Scalable Continual Reinforcement LearningMikel Malagón, Josu Ceberio, José Antonio LozanoICML 2024 · 被引用 13 次
- Stay Hungry, Keep Learning: Sustainable Plasticity for Deep Reinforcement LearningHuaicheng Zhou, Zifeng Zhuang, Donglin WangICML 2025
- Neuro-evolutionary Continual Reinforcement LearningPengyi Li, Hongyao Tang, Yifu Yuan, Yan Zheng 等ICML 2026
