Rethinking Momentum Knowledge Distillation in Online Continual Learning
Nicolas Michel, Maorong Wang, Ling Xiao, Toshihiko Yamasaki
摘要
Online Continual Learning (OCL) addresses the problem of training neural networks on a continuous data stream where multiple classification tasks emerge in sequence. In contrast to offline Continual Learning, data can be seen only once in OCL, which is a very severe constraint. In this context, replay-based strategies have achieved impressive results and most state-of-the-art approaches heavily depend on them. While Knowledge Distillation (KD) has been extensively used in offline Continual Learning, it remains under-exploited in OCL, despite its high potential. In this paper, we analyze the challenges in applying KD to OCL and give empirical justifications. We introduce a direct yet effective methodology for applying Momentum Knowledge Distillation (MKD) to many flagship OCL methods and demonstrate its capabilities to enhance existing approaches. In addition to improving existing state-of-the-art accuracy by more than points on ImageNet100, we shed light on MKD internal mechanics and impacts during training in OCL. We argue that similar to replay, MKD should be considered a central component of OCL. The code is available at https://github.com/Nicolas1203/mkd_ocl.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Model Sensitivity Aware Continual LearningZhenyi Wang, Heng HuangNeurIPS 2024 · 被引用 5 次
- Continual Distillation of Teachers from Different DomainsNicolas Michel, Maorong Wang, Jiangpeng He, Toshihiko YamasakiCVPR 2026 · 被引用 1 次
- PROL: Rehearsal Free Continual Learning in Streaming Data via Prompt Online LearningM. Anwar Ma'sum, Mahardhika Pratama, Savitha Ramasamy, Lin Liu 等ICCV 2025 · 被引用 1 次
- Rethinking the Stability-Plasticity Trade-off in Continual Learning from an Architectural PerspectiveAojun Lu, Hangjie Yuan, Tao Feng, Yanan SunICML 2025
- Perturbing to Preserve: Defending Fragile Knowledge in Online Continual LearningDulan Zhou, Zijian Gao, Kele XuAAAI 2026
它引用的顶会 Paper20
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna 等NeurIPS 2020 · 被引用 7,049 次
- Dark Experience for General Continual Learning: a Strong, Simple BaselinePietro Buzzega, Matteo Boschini, Angelo Porrello, Davide Abati 等NeurIPS 2020 · 被引用 1,494 次
- Contrastive Representation DistillationYonglong Tian, Dilip Krishnan, Phillip IsolaICLR 2020 · 被引用 1,305 次
- Decoupled Knowledge DistillationBorui Zhao, Quan Cui, Renjie Song, Yiyu Qiu 等CVPR 2022 · 被引用 835 次
相关 Paper
- Improving Plasticity in Online Continual Learning via Collaborative LearningMaorong Wang, Nicolas Michel, Ling Xiao, Toshihiko YamasakiCVPR 2024 · 被引用 7 次
- Online Prototype Learning for Online Continual LearningYujie Wei, Jiaxin Ye, Zhizhong Huang, Junping Zhang 等ICCV 2023 · 被引用 78 次
- Not Just Selection, but Exploration: Online Class-Incremental Continual Learning via Dual View ConsistencyYanan Gu, Xu Yang, Kun Wei, Cheng DengCVPR 2022 · 被引用 69 次
- Orchestrate Latent Expertise: Advancing Online Continual Learning with Multi-Level Supervision and Reverse Self-DistillationHongwei Yan, Liyuan Wang, Kaisheng Ma, Yi ZhongCVPR 2024 · 被引用 15 次
- Summarizing Stream Data for Memory-Constrained Online Continual LearningJianyang Gu, Kai Wang, Wei Jiang, Yang YouAAAI 2024 · 被引用 30 次
