Heads collapse, features stay: Why Replay needs big buffers
Giulia Lanzillotta, Damiano Meier, Thomas Hofmann
Abstract
A persistent paradox in continual learning (CL) is that neural networks often retain linearly separable representations of past tasks even when their output predictions fail. We formalize this distinction as the gap between deep (feature-space) and shallow (classifier-level) forgetting. We reveal a critical asymmetry in Experience Replay: while minimal buffers successfully anchor feature geometry and prevent deep forgetting, mitigating shallow forgetting typically requires substantially larger buffer capacities. To explain this, we extend the Neural Collapse framework to the sequential setting. We characterize deep forgetting as a geometric drift toward out-of-distribution subspaces and prove that any non-zero replay fraction asymptotically guarantees the retention of linear separability. Conversely, we identify that the ``strong collapse'' induced by small buffers leads to rank-deficient covariances and inflated class means, effectively blinding the classifier to true population boundaries. By unifying CL with out-of-distribution detection, our work challenges the prevailing reliance on large buffers, suggesting that explicitly correcting these statistical artifacts could unlock robust performance with minimal replay.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8c4b5056-6464-4545-b2c2-fcd3f562baf7Builds on16
- Dark Experience for General Continual Learning: a Strong, Simple BaselinePietro Buzzega, Matteo Boschini, Angelo Porrello, Davide Abati et al.NeurIPS 2020 · 1,494 citations
- Effect of scale on catastrophic forgetting in neural networksVinay Venkatesh Ramasesh, Aitor Lewkowycz, Ethan DyerICLR 2022 · 212 citations
- Anatomy of Catastrophic Forgetting: Hidden Representations and Task SemanticsVinay Venkatesh Ramasesh, Ethan Dyer, Maithra RaghuICLR 2021 · 207 citations
- Self-Supervised Models are Continual LearnersEnrico Fini, Victor G. Turrisi da Costa, Xavier Alameda-Pineda, Elisa Ricci et al.CVPR 2022 · 130 citations
- Extended Unconstrained Features Model for Exploring Deep Neural CollapseTom Tirer, Joan BrunaICML 2022 · 118 citations
Related papers
- Continual Out-of-Distribution Detection with Analytic Neural CollapseSaleh Momeni, Changnan Xiao, Bing LiuAAAI 2026
- Layerwise Proximal Replay: A Proximal Point Method for Online Continual LearningJinsoo Yoo, Yunpeng Liu, Frank Wood, Geoff PleissICML 2024 · 13 citations
- Towards Continual Learning Desiderata via HSIC-Bottleneck Orthogonalization and Equiangular EmbeddingDepeng Li, Tianqi Wang, Junwei Chen, Qining Ren et al.AAAI 2024 · 10 citations
- Predicting the Susceptibility of Examples to Catastrophic ForgettingGuy Hacohen, Tinne TuytelaarsICML 2025
- Retrospective Adversarial Replay for Continual LearningLilly Kumari, Shengjie Wang, Tianyi Zhou, Jeff A. BilmesNeurIPS 2022 · 57 citations
