MoCL: Metabolic Optimization for Curvature-Aware Continual Learning
Jiajun Lai, Qi Liu, Shijie Li, Huaiguang Jiang
Abstract
Continual learning requires models to mitigate catastrophic forgetting of prior knowledge while learning a sequence of tasks. Although existing methods based on orthogonal projection prevent interference by constraining parameter updates, they tend to limit plasticity as the task sequence progresses. The reliance on the linear approximation further causes the projected gradients to deviate from the nonlinear manifold. To address these issues, we propose Metabolic Optimization for Continual Learning (MoCL), a rehearsal-free framework that strikes a balance between stability and plasticity. To capture the geometric manifold of prior knowledge, MoCL introduces a factorized subspace approximation that avoids expensive explicit matrix inversion. Given the heavy-tailed distribution of the Fisher Information Matrix, we employ a metabolic gating based on Tsallis entropy to suppress updates that conflict with historical knowledge. Theoretical and empirical analyses show that MoCL suppresses interference while supporting shared low-loss behavior across sequential tasks. Extensive experimental results across multiple benchmarks demonstrate that MoCL outperforms state-of-the-art methods in both classification performance and efficiency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on26
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Moment Matching for Multi-Source Domain AdaptationXingchao Peng, Qinxun Bai, Xide Xia, Zijun Huang et al.ICCV 2019 · 2,239 citations
- Linear Mode Connectivity and the Lottery Ticket HypothesisJonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy, Michael CarbinICML 2020 · 750 citations
- Learning to Prompt for Continual LearningZifeng Wang, Zizhao Zhang, Chen-Yu Lee, Han Zhang et al.CVPR 2022 · 635 citations
- GaLore: Memory-Efficient LLM Training by Gradient Low-Rank ProjectionJiawei Zhao, Zhenyu Zhang, Beidi Chen, Zhangyang Wang et al.ICML 2024 · 433 citations
Related papers
- Octopus: History-Free Gradient Orthogonalization for Continual Learning in Multimodal Large Language ModelsYuehao Liu, Shanyan Guan, Weijia Zhang, Xuanming Shang et al.CVPR 2026
- Optimizing Spca-based Continual Learning: A Theoretical ApproachChunchun Yang, Malik Tiomoko, Zengfu WangICLR 2023
- Disentangling and mitigating the impact of task similarity for continual learningNaoki HirataniNeurIPS 2024 · 21 citations
- Continual Learning in Low-rank Orthogonal SubspacesArslan Chaudhry, Naeemullah Khan, Puneet K. Dokania, Philip H. S. TorrNeurIPS 2020 · 171 citations
- SplitLoRA: Balancing Stability and Plasticity in Continual Learning Through Gradient Space SplittingHaomiao Qiu, Miao Zhang, Ziyue Qiao, Weili Guan et al.ICLR 2026 · 10 citations
