Preserving Linear Separability in Continual Learning by Backward Feature Projection
Qiao Gu, Dongsub Shim, Florian Shkurti
Abstract
Catastrophic forgetting has been a major challenge in continual learning, where the model needs to learn new tasks with limited or no access to data from previously seen tasks. To tackle this challenge, methods based on knowledge distillation in feature space have been proposed and shown to reduce forgetting [16,19,27]. However, most feature distillation methods directly constrain the new features to match the old ones, overlooking the need for plasticity. To achieve a better stability-plasticity trade-off, we propose Backward Feature Projection (BFP), a method for continual learning that allows the new features to change up to a learnable linear transformation of the old features. BFP preserves the linear separability of the old classes while allowing the emergence of new feature directions to accommodate new classes. BFP can be integrated with existing experience replay methods and boost performance by a significant margin. We also demonstrate that BFP helps learn a better representation space, in which linear separability is well preserved during continual learning and linear probing achieves high classification accuracy. The code can be found at https://github.com/rvl-lab-utoronto/BFP.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3221c567-0b80-4c35-b70f-542592365505Cited by top-tier papers4
- IDER: IDempotent Experience Replay for Reliable Continual LearningZhanwang Liu, Yuting Li, Haoyuan Gao, Yexin Li et al.ICLR 2026 · 5 citations
- SEEKR: Selective Attention-Guided Knowledge Retention for Continual Learning of Large Language ModelsJinghan He, Haiyun Guo, Kuan Zhu, Zihan Zhao et al.EMNLP 2024 · 4 citations
- ESC: Erasing Space Concept for Knowledge DeletionTae-Young Lee, Sundong Park, Minwoo Jeon, Hyoseok Hwang et al.CVPR 2025
- DiAPR: Dimensionally-Allocated Prototype Refinement for Non-Exemplar Class Incremental LearningRuixuan Gao, Qijun Zhao, Keren FuAAAI 2026
Builds on14
- Co2L: Contrastive Continual LearningHyuntak Cha, Jaeho Lee, Jinwoo ShinICCV 2021 · 391 citations
- New Insights on Reducing Abrupt Representation Change in Online Continual LearningLucas Caccia, Rahaf Aljundi, Nader Asadi, Tinne Tuytelaars et al.ICLR 2022 · 279 citations
- Class-Incremental Learning via Dual AugmentationFei Zhu, Zhen Cheng, Xu-Yao Zhang, Cheng-Lin LiuNeurIPS 2021 · 256 citations
- Anatomy of Catastrophic Forgetting: Hidden Representations and Task SemanticsVinay Venkatesh Ramasesh, Ethan Dyer, Maithra RaghuICLR 2021 · 207 citations
- Towards Continual Knowledge Learning of Language ModelsJoel Jang, Seonghyeon Ye, Sohee Yang, Joongbo Shin et al.ICLR 2022 · 204 citations
Related papers
- Sketch-Based Replay Projection for Continual LearningJack Julian, Yun Sing Koh, Albert BifetKDD 2024 · 2 citations
- TRGP: Trust Region Gradient Projection for Continual LearningSen Lin, Li Yang, Deliang Fan, Junshan ZhangICLR 2022 · 107 citations
- Effective Continual Learning for Text Classification with Lightweight SnapshotsJue Wang, Dajie Dong, Lidan Shou, Ke Chen et al.AAAI 2023 · 4 citations
- Adaptive Plasticity Improvement for Continual LearningYan-Shuo Liang, Wu-Jun LiCVPR 2023
- Introducing Common Null Space of Gradients for Gradient Projection Methods in Continual LearningChengyi Yang, Mingda Dong, Xiaoyue Zhang, Jiayin Qi et al.ACM MM 2024 · 1 citation
