Orchestrate Latent Expertise: Advancing Online Continual Learning with Multi-Level Supervision and Reverse Self-Distillation
Hongwei Yan, Liyuan Wang, Kaisheng Ma, Yi Zhong
Abstract
To accommodate real-world dynamics, artificial intelligence systems need to cope with sequentially arriving content in an online manner. Beyond regular Continual Learning (CL) attempting to address catastrophic forgetting with offline training of each task, Online Continual Learning (OCL) is a more challenging yet realistic setting that performs CL in a one-pass data stream. Current OCL methods primarily rely on memory replay of old training samples. However, a notable gap from CL to OCL stems from the additional overfitting-underfitting dilemma associated with the use of rehearsal buffers: the inadequate learning of new training samples (underfitting) and the repeated learning of a few old training samples (overfitting). To this end, we introduce a novel approach, Multi-level Online Sequential Experts (MOSE), which cultivates the model as stacked sub-experts, integrating multi-level supervision and reverse self-distillation. Supervision signals across multiple stages facilitate appropriate convergence of the new task while gathering various strengths from experts by knowledge distillation mitigates the performance decline of old tasks. MOSE demonstrates remarkable efficacy in learning new samples and preserving past knowledge through multi-level experts, thereby significantly advancing OCL performance over state-of-the-art baselines (e.g., up to 7.3% on Split CIFAR-100 and 6.1% on Split Tiny-ImageNet).<sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">†</sup><sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">†</sup>Our code is available at https://github.com/AnAppleCore/MOSE
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 62e6fb86-9f2b-49af-a036-a36da6039c47Cited by top-tier papers14
- FlyPrompt: Brain-Inspired Random-Expanded Routing with Temporal-Ensemble Experts for General Continual LearningHongwei Yan, Guanglong Sun, Kanglei Zhou, Qian Li et al.ICLR 2026 · 4 citations
- MePo: Meta Post-Refinement for Rehearsal-Free General Continual LearningGuanglong Sun, Hongwei Yan, Liyuan Wang, Zhiqi KANG et al.ICML 2026 · 2 citations
- ConSurv: Multimodal Continual Learning for Survival AnalysisDianzhi Yu, Conghao Xiong, Yankai Chen, Wenqian Cui et al.AAAI 2026 · 2 citations
- VA-MoE: Variables-Adaptive Mixture of Experts for Incremental Weather ForecastingHao Chen, Tao Han, Song Guo, Jie Zhang et al.ICCV 2025 · 1 citation
- An Optimal Transport-driven Approach for Cultivating Latent Space in Online Incremental LearningQuyen Tran, Hai Nguyen, Minh Quan Dao, Hoang Phan et al.CVPR 2026
Builds on28
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Dark Experience for General Continual Learning: a Strong, Simple BaselinePietro Buzzega, Matteo Boschini, Angelo Porrello, Davide Abati et al.NeurIPS 2020 · 1,494 citations
- Be Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self DistillationLinfeng Zhang, Jiebo Song, Anni Gao, Jingwei Chen et al.ICCV 2019 · 1,069 citations
- Continual learning with hypernetworksJohannes von Oswald, Christian Henning, João Sacramento, Benjamin F. GreweICLR 2020 · 412 citations
- Co2L: Contrastive Continual LearningHyuntak Cha, Jaeho Lee, Jinwoo ShinICCV 2021 · 391 citations
Related papers
- Theory on Mixture-of-Experts in Continual LearningHongbo Li, Sen Lin, Lingjie Duan, Yingbin Liang et al.ICLR 2025
- Rethinking Momentum Knowledge Distillation in Online Continual LearningNicolas Michel, Maorong Wang, Ling Xiao, Toshihiko YamasakiICML 2024 · 26 citations
- Improving Plasticity in Online Continual Learning via Collaborative LearningMaorong Wang, Nicolas Michel, Ling Xiao, Toshihiko YamasakiCVPR 2024 · 7 citations
- PCR: Proxy-Based Contrastive Replay for Online Class-Incremental Continual LearningHuiwei Lin, Baoquan Zhang, Shanshan Feng, Xutao Li et al.CVPR 2023
- Progressive Prototype Evolving for Dual-Forgetting Mitigation in Non-Exemplar Online Continual LearningQiwei Li, Yuxin Peng, Jiahuan ZhouACM MM 2024 · 4 citations
