Rethinking Momentum Knowledge Distillation in Online Continual Learning
Nicolas Michel, Maorong Wang, Ling Xiao, Toshihiko Yamasaki
Abstract
Online Continual Learning (OCL) addresses the problem of training neural networks on a continuous data stream where multiple classification tasks emerge in sequence. In contrast to offline Continual Learning, data can be seen only once in OCL, which is a very severe constraint. In this context, replay-based strategies have achieved impressive results and most state-of-the-art approaches heavily depend on them. While Knowledge Distillation (KD) has been extensively used in offline Continual Learning, it remains under-exploited in OCL, despite its high potential. In this paper, we analyze the challenges in applying KD to OCL and give empirical justifications. We introduce a direct yet effective methodology for applying Momentum Knowledge Distillation (MKD) to many flagship OCL methods and demonstrate its capabilities to enhance existing approaches. In addition to improving existing state-of-the-art accuracy by more than points on ImageNet100, we shed light on MKD internal mechanics and impacts during training in OCL. We argue that similar to replay, MKD should be considered a central component of OCL. The code is available at https://github.com/Nicolas1203/mkd_ocl.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 70ebc87a-bd4d-442b-94a8-0c9367b0ad09Cited by top-tier papers7
- Model Sensitivity Aware Continual LearningZhenyi Wang, Heng HuangNeurIPS 2024 · 5 citations
- Continual Distillation of Teachers from Different DomainsNicolas Michel, Maorong Wang, Jiangpeng He, Toshihiko YamasakiCVPR 2026 · 1 citation
- PROL: Rehearsal Free Continual Learning in Streaming Data via Prompt Online LearningM. Anwar Ma'sum, Mahardhika Pratama, Savitha Ramasamy, Lin Liu et al.ICCV 2025 · 1 citation
- Rethinking the Stability-Plasticity Trade-off in Continual Learning from an Architectural PerspectiveAojun Lu, Hangjie Yuan, Tao Feng, Yanan SunICML 2025
- Perturbing to Preserve: Defending Fragile Knowledge in Online Continual LearningDulan Zhou, Zijian Gao, Kele XuAAAI 2026
Builds on20
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna et al.NeurIPS 2020 · 7,049 citations
- Dark Experience for General Continual Learning: a Strong, Simple BaselinePietro Buzzega, Matteo Boschini, Angelo Porrello, Davide Abati et al.NeurIPS 2020 · 1,494 citations
- Contrastive Representation DistillationYonglong Tian, Dilip Krishnan, Phillip IsolaICLR 2020 · 1,305 citations
- Decoupled Knowledge DistillationBorui Zhao, Quan Cui, Renjie Song, Yiyu Qiu et al.CVPR 2022 · 835 citations
Related papers
- Improving Plasticity in Online Continual Learning via Collaborative LearningMaorong Wang, Nicolas Michel, Ling Xiao, Toshihiko YamasakiCVPR 2024 · 7 citations
- Online Prototype Learning for Online Continual LearningYujie Wei, Jiaxin Ye, Zhizhong Huang, Junping Zhang et al.ICCV 2023 · 78 citations
- Not Just Selection, but Exploration: Online Class-Incremental Continual Learning via Dual View ConsistencyYanan Gu, Xu Yang, Kun Wei, Cheng DengCVPR 2022 · 69 citations
- Orchestrate Latent Expertise: Advancing Online Continual Learning with Multi-Level Supervision and Reverse Self-DistillationHongwei Yan, Liyuan Wang, Kaisheng Ma, Yi ZhongCVPR 2024 · 15 citations
- Summarizing Stream Data for Memory-Constrained Online Continual LearningJianyang Gu, Kai Wang, Wei Jiang, Yang YouAAAI 2024 · 30 citations
