Class Incremental Learning with Multi-Teacher Distillation
Haitao Wen, Lili Pan, Yu Dai, Heqian Qiu, Lanxiao Wang, Qingbo Wu, Hongliang Li
Abstract
Distillation strategies are currently the primary approaches for mitigating forgetting in class incremental learning (CIL). Existing methods generally inherit previous knowledge from a single teacher. However, teachers with different mechanisms are talented at different tasks, and inheriting diverse knowledge from them can enhance compatibility with new knowledge. In this paper, we propose the MTD method to find multiple diverse teachers for CIL. Specifically, we adopt weight permutation, feature perturbation, and diversity regularization techniques to ensure diverse mechanisms in teachers. To reduce time and memory consumption, each teacher is represented as a small branch in the model. We adapt existing CIL distillation strategies with MTD and extensive experiments on CIFAR-100, ImageNet-100, and ImageNet-1000 show significant performance improvement. Our code is available at https://github.com/HaitaoWen/CLearning .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3da2c823-56d1-4265-83fb-d734f82785c0Cited by top-tier papers14
- Mixture of Noise for Pre-Trained Model-Based Class-Incremental LearningKai Jiang, Zhengyan Shi, Dell Zhang, Hongyuan Zhang et al.NeurIPS 2025 · 38 citations
- Continual Gaussian Mixture Distribution Modeling for Class Incremental Semantic SegmentationGuilin Zhu, Runmin Wang, Yuanjie Shao, Weidong Yang et al.NeurIPS 2025 · 5 citations
- Overcoming Dual Drift for Continual Long-Tailed Visual Question AnsweringFeifei Zhang, Zhihao Wang, Xi Zhang, Changsheng XuICCV 2025 · 3 citations
- Scaling Continual Learning to 300+ Tasks with Bi-Level Routing Mixture-of-ExpertsMeng Lou, Yunxiang Fu, Yizhou YuICML 2026 · 1 citation
- Swiss Army Knife: Synergizing Biases in Knowledge from Vision Foundation Models for Multi-Task LearningYuxiang Lu, Shengcao Cao, Yu-Xiong WangICLR 2025
Builds on19
- Gradient Projection Memory for Continual LearningGobinda Saha, Isha Garg, Kaushik RoyICLR 2021 · 409 citations
- The Role of Permutation Invariance in Linear Mode Connectivity of Neural NetworksRahim Entezari, Hanie Sedghi, Olga Saukh, Behnam NeyshaburICLR 2022 · 301 citations
- SS-IL: Separated Softmax for Incremental LearningHongjoon Ahn, Jihwan Kwak, Subin Lim, Hyeonsu Bang et al.ICCV 2021 · 209 citations
- Class-Incremental Learning by Knowledge Distillation with Adaptive Feature ConsolidationMinsoo Kang, Jaeyoo Park, Bohyung HanCVPR 2022 · 189 citations
- Geometry of the Loss Landscape in Overparameterized Neural Networks: Symmetries and InvariancesBerfin Simsek, François Ged, Arthur Jacot, Francesco Spadaro et al.ICML 2021 · 136 citations
Related papers
- Few-Shot Class-Incremental Learning via Class-Aware Bilateral DistillationLinglan Zhao, Jing Lu, Yunlu Xu, Zhanzhan Cheng et al.CVPR 2023
- M2SD: Multiple Mixing Self-Distillation for Few-Shot Class-Incremental LearningJinhao Lin, Ziheng Wu, Weifeng Lin, Jun Huang et al.AAAI 2024 · 13 citations
- Distilling Causal Effect of Data in Class-Incremental LearningXinting Hu, Kaihua Tang, Chunyan Miao, Xian-Sheng Hua et al.CVPR 2021
- Resolving Task Confusion in Dynamic Expansion Architectures for Class Incremental LearningBingchen Huang, Zhineng Chen, Peng Zhou, Jiayin Chen et al.AAAI 2023 · 31 citations
- Agree to Disagree: Adaptive Ensemble Knowledge Distillation in Gradient SpaceShangchen Du, Shan You, Xiaojie Li, Jianlong Wu et al.NeurIPS 2020 · 144 citations
