Beyond Myopic Alignment: Lookahead Optimization for Online Class-Incremental Learning
Song Lai, Zhe Zhao, Fei Zhu, Ji Cheng, Xi Lin, Qingfu Zhang, Gaofeng Meng
Abstract
Rehearsal-based methods are the cornerstone of modern online class-incremental learning (OCIL), yet they face a fundamental challenge: the gradient of the current task often conflicts with that of the rehearsal data from the memory buffer, leading to catastrophic forgetting. Recent works have implicitly addressed this by using hypergradients, but the underlying mechanism has remained poorly understood. In this paper, we provide a formal analysis revealing that hypergradients mitigate forgetting by aligning task-specific gradients towards a common meta-objective, thereby reducing their conflict. However, we argue that this conflictreducing alignment is inherently myopic-it only considers the immediate gradient directions, failing to account for the loss landscape geometry one step ahead. To overcome this limitation, we introduce a novel framework: Lookahead Optimization for Rehearsal (LOR). LOR explores a set of future model states by taking lookahead steps along different directions that balance plasticity and stability and optimizes a first-order Log-Sum-Exp (LSE) surrogate to emphasize the worst-performing sampled lookahead directions. Theoretical analysis from both optimization and statistical perspectives corroborates the robustness of our approach. Extensive experiments on Seq-CIFAR10, Seq-CIFAR100, and Seq-TinyImageNet demonstrate that LOR significantly outperforms state-of-the-art methods, establishing a new and more robust paradigm for rehearsal-based OCIL.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on16
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine et al.NeurIPS 2020 · 2,261 citations
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 1,861 citations
- Distributionally Robust Neural NetworksShiori Sagawa, Pang Wei Koh, Tatsunori B. Hashimoto, Percy LiangICLR 2020 · 1,578 citations
- Dark Experience for General Continual Learning: a Strong, Simple BaselinePietro Buzzega, Matteo Boschini, Angelo Porrello, Davide Abati et al.NeurIPS 2020 · 1,494 citations
- Conflict-Averse Gradient Descent for Multi-task learningBo Liu, Xingchao Liu, Xiaojie Jin, Peter Stone et al.NeurIPS 2021 · 686 citations
Related papers
- Mitigating Catastrophic Forgetting in Online Continual Learning by Modeling Previous Task Interrelations via Pareto OptimizationYichen Wu, Hong Wang, Peilin Zhao, Yefeng Zheng et al.ICML 2024 · 23 citations
- Gradient-Guided Epsilon Constraint Method for Online Continual LearningSong Lai, Changyi Ma, Fei Zhu, Zhe Zhao et al.NeurIPS 2025 · 3 citations
- Heterogeneous Forgetting Compensation for Class-Incremental LearningJiahua Dong, Wenqi Liang, Yang Cong, Gan SunICCV 2023 · 28 citations
- Unlocking the Power of Rehearsal in Continual Learning: A Theoretical PerspectiveJunze Deng, Qinhang Wu, Peizhong Ju, Sen Lin et al.ICML 2025
- Retrospective Adversarial Replay for Continual LearningLilly Kumari, Shengjie Wang, Tianyi Zhou, Jeff A. BilmesNeurIPS 2022 · 57 citations
