Meta Continual Learning Revisited: Implicitly Enhancing Online Hessian Approximation via Variance Reduction
Yichen Wu, Long-Kai Huang, Renzhen Wang, Deyu Meng, Ying Wei
摘要
Regularization-based methods have so far been among the de facto choices for continual learning. Recent theoretical studies have revealed that these methods all boil down to relying on the Hessian matrix approximation of model weights. However, these methods suffer from suboptimal trade-offs between knowledge transfer and forgetting due to fixed and unchanging Hessian estimations during training. Another seemingly parallel strand of Meta-Continual Learning (Meta-CL) algorithms enforces alignment between gradients of previous tasks and that of the current task. In this work we revisit Meta-CL and for the first time bridge it with regularization-based methods. Concretely, Meta-CL implicitly approximates Hessian in an online manner, which enjoys the benefits of timely adaptation but meantime suffers from high variance induced by random memory buffer sampling. We are thus highly motivated to combine the best of both worlds, through the proposal of Variance Reduced Meta-CL (VR-MCL) to achieve both timely and accurate Hessian approximation. Through comprehensive experiments across three datasets and various settings, we consistently observe that VR-MCL outperforms other SOTA methods, which further validates the effectiveness of VR-MCL 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- DuQuant: Distributing Outliers via Dual Transformation Makes Stronger Quantized LLMsHaokun Lin, Haobo Xu, Yichen Wu, Jingzhi Cui 等NeurIPS 2024 · 被引用 206 次
- Unified Gradient-Based Machine Unlearning with Remain Geometry EnhancementZhehao Huang, Xinwen Cheng, JingHao Zheng, Haoran Wang 等NeurIPS 2024 · 被引用 57 次
- Mitigating Catastrophic Forgetting in Online Continual Learning by Modeling Previous Task Interrelations via Pareto OptimizationYichen Wu, Hong Wang, Peilin Zhao, Yefeng Zheng 等ICML 2024 · 被引用 23 次
- Federated Continual Learning via Prompt-based Dual Knowledge TransferHongming Piao, Yichen Wu, Dapeng Wu, Ying WeiICML 2024 · 被引用 19 次
- MINGLE: Mixture of Null-Space Gated Low-Rank Experts for Test-Time Continual Model MergingZihuan Qiu, Yi Xu, Chiyuan He, Fanman Meng 等NeurIPS 2025 · 被引用 16 次
它引用的顶会 Paper21
- Dark Experience for General Continual Learning: a Strong, Simple BaselinePietro Buzzega, Matteo Boschini, Angelo Porrello, Davide Abati 等NeurIPS 2020 · 被引用 1,494 次
- Coresets via Bilevel Optimization for Continual Learning and StreamingZalán Borsos, Mojmir Mutny, Andreas KrauseNeurIPS 2020 · 被引用 320 次
- New Insights on Reducing Abrupt Representation Change in Online Continual LearningLucas Caccia, Rahaf Aljundi, Nader Asadi, Tinne Tuytelaars 等ICLR 2022 · 被引用 279 次
- Online Class-Incremental Continual Learning with Adversarial Shapley ValueDongsub Shim, Zheda Mai, Jihwan Jeong, Scott Sanner 等AAAI 2021 · 被引用 262 次
- A Near-Optimal Algorithm for Stochastic Bilevel Optimization via Double-MomentumPrashant Khanduri, Siliang Zeng, Mingyi Hong, Hoi-To Wai 等NeurIPS 2021 · 被引用 175 次
相关 Paper
- Towards Dynamic Modality Alignment in Multimodal Continual LearningJiayao Tan, Fan Lyu, Tianle Liu, Fuyuan Hu 等CVPR 2026
- Memory-Statistics Tradeoff in Continual Learning with Structural RegularizationHaoran Li, Jingfeng Wu, Vladimir BravermanICLR 2026 · 被引用 4 次
- An Effective Dynamic Gradient Calibration Method for Continual LearningWeichen Lin, Jiaxiang Chen, Ruomin Huang, Hu DingICML 2024 · 被引用 2 次
- DualHSIC: HSIC-Bottleneck and Alignment for Continual LearningZifeng Wang, Zheng Zhan, Yifan Gong, Yucai Shao 等ICML 2023 · 被引用 21 次
- The Ideal Continual Learner: An Agent That Never ForgetsLiangzu Peng, Paris Giampouras, René VidalICML 2023 · 被引用 39 次
