Beyond Sharpness: A Flatness Decomposition Framework for Efficient Continual Learning
Yanan Chen, Tieliang Gong, Yunjiao Zhang, Wen Wen
摘要
Continual Learning (CL) aims to enable models to sequentially learn multiple tasks without forgetting previous knowledge. Recent studies have shown that optimizing towards flatter loss minima can improve model generalization. However, existing sharpness-aware methods for CL suffer from two key limitations: (1) they treat sharpness regularization as a unified signal without distinguishing the contributions of its components. and (2) they introduce substantial computational overhead that impedes practical deployment. To address these challenges, we propose FLAD, a novel optimization framework that decomposes sharpness-aware perturbations into gradient-aligned and stochastic-noise components, and show that retaining only the noise component promotes generalization. We further introduce a lightweight scheduling scheme that enables FLAD to maintain significant performance gains even under constrained training time. FLAD can be seamlessly integrated into various CL paradigms and consistently outperforms standard and sharpness-aware optimizers in diverse experimental settings, demonstrating its effectiveness and practicality in CL.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper17
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 被引用 1,861 次
- Gradient Projection Memory for Continual LearningGobinda Saha, Isha Garg, Kaushik RoyICLR 2021 · 被引用 409 次
- Overcoming Catastrophic Forgetting in Incremental Few-Shot Learning by Finding Flat MinimaGuangyuan Shi, Jiaxin Chen, Wenlong Zhang, Li-Ming Zhan 等NeurIPS 2021 · 被引用 229 次
- Towards Understanding Sharpness-Aware MinimizationMaksym Andriushchenko, Nicolas FlammarionICML 2022 · 被引用 190 次
- Flattening Sharpness for Dynamic Gradient Projection Memory Benefits Continual LearningDanruo Deng, Guangyong Chen, Jianye Hao, Qiong Wang 等NeurIPS 2021 · 被引用 112 次
相关 Paper
- Make Continual Learning Stronger via C-FlatAng Bian, Wei Li, Hangjie Yuan, Chengrong Yu 等NeurIPS 2024 · 被引用 48 次
- A Faster Path to Continual LearningWei Li, Hangjie Yuan, Zixiang Zhao, Borui Kang 等CVPR 2026 · 被引用 2 次
- Lookbehind-SAM: k steps back, 1 step forwardGonçalo Mordido, Pranshu Malviya, Aristide Baratin, Sarath ChandarICML 2024 · 被引用 4 次
- Data Augmented Flatness-aware Gradient Projection for Continual LearningEnneng Yang, Li Shen, Zhenyi Wang, Shiwei Liu 等ICCV 2023 · 被引用 28 次
- Continual Learners are Incremental Model GeneralizersJaehong Yoon, Sung Ju Hwang, Yue CaoICML 2023 · 被引用 6 次
