Make Continual Learning Stronger via C-Flat
Ang Bian, Wei Li, Hangjie Yuan, Chengrong Yu, Mang Wang, Zixiang Zhao, Aojun Lu, Pengliang Ji, Tao Feng
Abstract
How to balance the learning 'sensitivity-stability' upon new task training and memory preserving is critical in CL to resolve catastrophic forgetting. Improving model generalization ability within each learning phase is one solution to help CL learning overcome the gap in the joint knowledge space. Zeroth-order loss landscape sharpness-aware minimization is a strong training regime improving model generalization in transfer learning compared with optimizer like SGD. It has also been introduced into CL to improve memory representation or learning efficiency. However, zeroth-order sharpness alone could favors sharper over flatter minima in certain scenarios, leading to a rather sensitive minima rather than a global optima. To further enhance learning stability, we propose a Continual Flatness (C-Flat) method featuring a flatter loss landscape tailored for CL. C-Flat could be easily called with only one line of code and is plug-and-play to any CL methods. A general framework of C-Flat applied to all CL categories and a thorough comparison with loss minima optimizer and flat minima based CL approaches is presented in this paper, showing that our method can boost CL performance in almost all cases.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4de5308e-d75e-4f5e-8235-15d068781badCited by top-tier papers19
- Continual Multimodal Contrastive LearningXiaohao Liu, Xiaobo Xia, See-Kiong Ng, Tat-Seng ChuaNeurIPS 2025 · 25 citations
- MathFlow: Enhancing the Perceptual Flow of MLLMs for Visual Mathematical ProblemsShuhang Chen, Hangjie Yuan, Yunqiu Xu, Pengwei Liu et al.ACL 2026 · 9 citations
- Sharpness-Aware Pretraining Mitigates Catastrophic ForgettingIshaan Watts, Catherine Li, Sachin Goyal, Jacob Mitchell Springer et al.ICML 2026 · 6 citations
- Continual GUI AgentsZiwei Liu, Borui Kang, Hangjie Yuan, Zixiang Zhao et al.ICML 2026 · 6 citations
- Dynamic Multi-Layer Null Space Projection for Vision-Language Continual LearningBorui Kang, Lei Wang, Zhiping Wu, Tao Feng et al.ICCV 2025 · 4 citations
Builds on35
- PaLM-E: An Embodied Multimodal Language ModelDanny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch et al.ICML 2023 · 2,601 citations
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 1,861 citations
- Gradient Projection Memory for Continual LearningGobinda Saha, Isha Garg, Kaushik RoyICLR 2021 · 409 citations
- Gradient Vaccine: Investigating and Improving Multi-task Optimization in Massively Multilingual ModelsZirui Wang, Yulia Tsvetkov, Orhan Firat, Yuan CaoICLR 2021 · 241 citations
- Overcoming Catastrophic Forgetting in Incremental Few-Shot Learning by Finding Flat MinimaGuangyuan Shi, Jiaxin Chen, Wenlong Zhang, Li-Ming Zhan et al.NeurIPS 2021 · 229 citations
Related papers
- Beyond Sharpness: A Flatness Decomposition Framework for Efficient Continual LearningYanan Chen, Tieliang Gong, Yunjiao Zhang, Wen WenAAAI 2026
- A Faster Path to Continual LearningWei Li, Hangjie Yuan, Zixiang Zhao, Borui Kang et al.CVPR 2026 · 2 citations
- Flattening Sharpness for Dynamic Gradient Projection Memory Benefits Continual LearningDanruo Deng, Guangyong Chen, Jianye Hao, Qiong Wang et al.NeurIPS 2021 · 112 citations
- Data Augmented Flatness-aware Gradient Projection for Continual LearningEnneng Yang, Li Shen, Zhenyi Wang, Shiwei Liu et al.ICCV 2023 · 28 citations
- Random Amalgamation of Adapters for Flatter Loss Landscapes: Towards Class-Incremental Learning with Better StabilityYao Deng, Xiang Xiang, Jiaqi GuiAAAI 2026
