Meta Continual Learning Revisited: Implicitly Enhancing Online Hessian Approximation via Variance Reduction
Yichen Wu, Long-Kai Huang, Renzhen Wang, Deyu Meng, Ying Wei
Abstract
Regularization-based methods have so far been among the de facto choices for continual learning. Recent theoretical studies have revealed that these methods all boil down to relying on the Hessian matrix approximation of model weights. However, these methods suffer from suboptimal trade-offs between knowledge transfer and forgetting due to fixed and unchanging Hessian estimations during training. Another seemingly parallel strand of Meta-Continual Learning (Meta-CL) algorithms enforces alignment between gradients of previous tasks and that of the current task. In this work we revisit Meta-CL and for the first time bridge it with regularization-based methods. Concretely, Meta-CL implicitly approximates Hessian in an online manner, which enjoys the benefits of timely adaptation but meantime suffers from high variance induced by random memory buffer sampling. We are thus highly motivated to combine the best of both worlds, through the proposal of Variance Reduced Meta-CL (VR-MCL) to achieve both timely and accurate Hessian approximation. Through comprehensive experiments across three datasets and various settings, we consistently observe that VR-MCL outperforms other SOTA methods, which further validates the effectiveness of VR-MCL 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers21
- DuQuant: Distributing Outliers via Dual Transformation Makes Stronger Quantized LLMsHaokun Lin, Haobo Xu, Yichen Wu, Jingzhi Cui et al.NeurIPS 2024 · 206 citations
- Unified Gradient-Based Machine Unlearning with Remain Geometry EnhancementZhehao Huang, Xinwen Cheng, JingHao Zheng, Haoran Wang et al.NeurIPS 2024 · 57 citations
- Mitigating Catastrophic Forgetting in Online Continual Learning by Modeling Previous Task Interrelations via Pareto OptimizationYichen Wu, Hong Wang, Peilin Zhao, Yefeng Zheng et al.ICML 2024 · 23 citations
- Federated Continual Learning via Prompt-based Dual Knowledge TransferHongming Piao, Yichen Wu, Dapeng Wu, Ying WeiICML 2024 · 19 citations
- MINGLE: Mixture of Null-Space Gated Low-Rank Experts for Test-Time Continual Model MergingZihuan Qiu, Yi Xu, Chiyuan He, Fanman Meng et al.NeurIPS 2025 · 16 citations
Builds on21
- Dark Experience for General Continual Learning: a Strong, Simple BaselinePietro Buzzega, Matteo Boschini, Angelo Porrello, Davide Abati et al.NeurIPS 2020 · 1,494 citations
- Coresets via Bilevel Optimization for Continual Learning and StreamingZalán Borsos, Mojmir Mutny, Andreas KrauseNeurIPS 2020 · 320 citations
- New Insights on Reducing Abrupt Representation Change in Online Continual LearningLucas Caccia, Rahaf Aljundi, Nader Asadi, Tinne Tuytelaars et al.ICLR 2022 · 279 citations
- Online Class-Incremental Continual Learning with Adversarial Shapley ValueDongsub Shim, Zheda Mai, Jihwan Jeong, Scott Sanner et al.AAAI 2021 · 262 citations
- A Near-Optimal Algorithm for Stochastic Bilevel Optimization via Double-MomentumPrashant Khanduri, Siliang Zeng, Mingyi Hong, Hoi-To Wai et al.NeurIPS 2021 · 175 citations
Related papers
- Towards Dynamic Modality Alignment in Multimodal Continual LearningJiayao Tan, Fan Lyu, Tianle Liu, Fuyuan Hu et al.CVPR 2026
- Memory-Statistics Tradeoff in Continual Learning with Structural RegularizationHaoran Li, Jingfeng Wu, Vladimir BravermanICLR 2026 · 4 citations
- An Effective Dynamic Gradient Calibration Method for Continual LearningWeichen Lin, Jiaxiang Chen, Ruomin Huang, Hu DingICML 2024 · 2 citations
- DualHSIC: HSIC-Bottleneck and Alignment for Continual LearningZifeng Wang, Zheng Zhan, Yifan Gong, Yucai Shao et al.ICML 2023 · 21 citations
- The Ideal Continual Learner: An Agent That Never ForgetsLiangzu Peng, Paris Giampouras, René VidalICML 2023 · 39 citations
