FOREVER: Forgetting Curve-Inspired Memory Replay for Language Model Continual Learning
Yujie Feng, Hao Wang, Jian Li, Xu Chu, Zhaolu Kang, Yiran Liu, Yasha Wang, Philip S. Yu, Xiao-Ming Wu
Abstract
Continual learning (CL) for large language models (LLMs) aims to enable sequential knowledge acquisition without catastrophic forgetting. Memory replay methods are widely used for their practicality and effectiveness, but most rely on fixed, step-based heuristics that often misalign with the model's actual learning progress, since identical training steps can result in varying degrees of parameter change. Motivated by recent findings that LLM forgetting mirrors the Ebbinghaus human forgetting curve, we propose FOREVER (FORgEtting curVe-inspired mEmory Replay), a novel CL framework that aligns replay schedules with a model-centric notion of time. FOREVER defines model time using the magnitude of optimizer updates, allowing forgetting curveinspired replay intervals to align with the model's internal evolution rather than raw training steps. Building on this approach, FOR-EVER incorporates a forgetting curve-based replay scheduler to determine when to replay and an intensity-aware regularization mechanism to adaptively control how to replay. Extensive experiments on three CL benchmarks and models ranging from 0.6B to 13B parameters demonstrate that FOREVER consistently mitigates catastrophic forgetting 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6e3eed58-bb6a-4190-b738-4d306f199d08Cited by top-tier papers2
- Micro-Macro Retrieval: Reducing Long-Form Hallucination in Large Language ModelsYujie Feng, Jian Li, Zhihan Zhou, Pengfei Xu et al.ICLR 2026
- Lightweight Federated Incremental Learning via Decoupled ReplayXiuying Wang, Yichen Li, Hang Su, Gaozhuo Liu et al.ICML 2026
Builds on29
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- MemoryBank: Enhancing Large Language Models with Long-Term MemoryWanjun Zhong, Lianghong Guo, Qiqi Gao, He Ye et al.AAAI 2024 · 394 citations
- Achieving Forgetting Prevention and Knowledge Transfer in Continual LearningZixuan Ke, Bing Liu, Nianzu Ma, Hu Xu et al.NeurIPS 2021 · 167 citations
- Knowledge Fusion of Large Language ModelsFanqi Wan, Xinting Huang, Deng Cai, Xiaojun Quan et al.ICLR 2024 · 113 citations
- Entity Alignment with Noisy Annotations from Large Language ModelsShengyuan Chen, Qinggang Zhang, Junnan Dong, Wen Hua et al.NeurIPS 2024 · 44 citations
Related papers
- SEEKR: Selective Attention-Guided Knowledge Retention for Continual Learning of Large Language ModelsJinghan He, Haiyun Guo, Kuan Zhu, Zihan Zhao et al.EMNLP 2024 · 4 citations
- Do Your Best and Get Enough Rest for Continual LearningHankyul Kang, Gregor Seifer, Donghyun Lee, Jongbin RyuCVPR 2025
- Progressive Prompts: Continual Learning for Language ModelsAnastasia Razdaibiedina, Yuning Mao, Rui Hou, Madian Khabsa et al.ICLR 2023 · 15 citations
- Recurrent Knowledge Identification and Fusion for Language Model Continual LearningYujie Feng, Xujia Wang, Zexin Lu, Shenghong Fu et al.ACL 2025
- Self-Evolving Pseudo-Rehearsal for Catastrophic Forgetting with Task Similarity in LLMsJun Wang, Liang Ding, Shuai Wang, Hongyu Li et al.NeurIPS 2025 · 4 citations
