Beyond Sharpness: A Flatness Decomposition Framework for Efficient Continual Learning
Yanan Chen, Tieliang Gong, Yunjiao Zhang, Wen Wen
Abstract
Continual Learning (CL) aims to enable models to sequentially learn multiple tasks without forgetting previous knowledge. Recent studies have shown that optimizing towards flatter loss minima can improve model generalization. However, existing sharpness-aware methods for CL suffer from two key limitations: (1) they treat sharpness regularization as a unified signal without distinguishing the contributions of its components. and (2) they introduce substantial computational overhead that impedes practical deployment. To address these challenges, we propose FLAD, a novel optimization framework that decomposes sharpness-aware perturbations into gradient-aligned and stochastic-noise components, and show that retaining only the noise component promotes generalization. We further introduce a lightweight scheduling scheme that enables FLAD to maintain significant performance gains even under constrained training time. FLAD can be seamlessly integrated into various CL paradigms and consistently outperforms standard and sharpness-aware optimizers in diverse experimental settings, demonstrating its effectiveness and practicality in CL.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e3d527d7-8159-442f-8fa5-84e90c2a996cBuilds on17
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 1,861 citations
- Gradient Projection Memory for Continual LearningGobinda Saha, Isha Garg, Kaushik RoyICLR 2021 · 409 citations
- Overcoming Catastrophic Forgetting in Incremental Few-Shot Learning by Finding Flat MinimaGuangyuan Shi, Jiaxin Chen, Wenlong Zhang, Li-Ming Zhan et al.NeurIPS 2021 · 229 citations
- Towards Understanding Sharpness-Aware MinimizationMaksym Andriushchenko, Nicolas FlammarionICML 2022 · 190 citations
- Flattening Sharpness for Dynamic Gradient Projection Memory Benefits Continual LearningDanruo Deng, Guangyong Chen, Jianye Hao, Qiong Wang et al.NeurIPS 2021 · 112 citations
Related papers
- Make Continual Learning Stronger via C-FlatAng Bian, Wei Li, Hangjie Yuan, Chengrong Yu et al.NeurIPS 2024 · 48 citations
- A Faster Path to Continual LearningWei Li, Hangjie Yuan, Zixiang Zhao, Borui Kang et al.CVPR 2026 · 2 citations
- Lookbehind-SAM: k steps back, 1 step forwardGonçalo Mordido, Pranshu Malviya, Aristide Baratin, Sarath ChandarICML 2024 · 4 citations
- Data Augmented Flatness-aware Gradient Projection for Continual LearningEnneng Yang, Li Shen, Zhenyi Wang, Shiwei Liu et al.ICCV 2023 · 28 citations
- Continual Learners are Incremental Model GeneralizersJaehong Yoon, Sung Ju Hwang, Yue CaoICML 2023 · 6 citations
