Reawakening knowledge: Anticipatory recovery from catastrophic interference via structured training
Yanlai Yang, Matt Jones, Michael C. Mozer, Mengye Ren
Abstract
We explore the training dynamics of neural networks in a structured non-IID setting where documents are presented cyclically in a fixed, repeated sequence. Typically, networks suffer from catastrophic interference when training on a sequence of documents; however, we discover a curious and remarkable property of LLMs finetuned sequentially in this setting: they exhibit anticipatory behavior, recovering from the forgetting on documents before encountering them again. This behavior occurs even though the documents are never presented in context together. The behavior emerges and becomes more robust as the architecture scales up its number of parameters. Through comprehensive experiments and visualizations, we demonstrate a new mechanism by which over-parametrized neural networks can recover from catastrophic interference and uncover new insights into training over-parameterized networks in cyclically structured environments.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b4f7a868-23ab-4142-9dc0-ac868b9cfa1eCited by top-tier papers3
- From Dormant to Deleted: Tamper-Resistant Unlearning Through Weight-Space RegularizationShoaib Ahmed Siddiqui, Adrian Weller, David Krueger, Gintare Karolina Dziugaite et al.NeurIPS 2025 · 16 citations
- Conformal Arbitrage: Risk-Controlled Balancing of Competing Objectives in Language ModelsWilliam Overman, Mohsen BayatiNeurIPS 2025 · 12 citations
- Fresh in memory: Training-order recency is linearly encoded in language model activationsDmitrii Krasheninnikov, Richard E. Turner, David KruegerICLR 2026 · 5 citations
Builds on24
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
- Generative Pretraining From PixelsMark Chen, Alec Radford, Rewon Child, Jeffrey Wu et al.ICML 2020 · 1,773 citations
Related papers
- Organizing recurrent network dynamics by task-computation to enable continual learningLea Duncker, Laura Driscoll, Krishna V. Shenoy, Maneesh Sahani et al.NeurIPS 2020 · 108 citations
- Provable Long-Range Benefits of Next-Token PredictionXinyuan Cao, Santosh S. VempalaSTOC 2026
- Theory on Forgetting and Generalization of Continual LearningSen Lin, Peizhong Ju, Yingbin Liang, Ness B. ShroffICML 2023 · 74 citations
- Rehearsal revealed: The limits and merits of revisiting samples in continual learningEli Verwimp, Matthias De Lange, Tinne TuytelaarsICCV 2021 · 121 citations
- The Importance of Being Lazy: Scaling Limits of Continual LearningJacopo Graldi, Alessandro Breccia, Giulia Lanzillotta, Thomas Hofmann et al.ICML 2025
