Measuring Asymmetric Gradient Discrepancy in Parallel Continual Learning
Fan Lyu, Qing Sun, Fanhua Shang, Liang Wan, Wei Feng
Abstract
In Parallel Continual Learning (PCL), the parallel multiple tasks start and end training unpredictably, thus suffering from both training conflict and catastrophic forgetting issues. The two issues are raised because the gradients from parallel tasks differ in directions and magnitudes. Thus, in this paper, we formulate the PCL into a minimum distance optimization problem among gradients and propose an explicit Asymmetric Gradient Distance (AGD) to evaluate the gradient discrepancy in PCL. AGD considers both gradient magnitude ratios and directions, and has a tolerance when updating with a small gradient of inverse direction, which reduces the imbalanced influence of gradients on parallel task training. Moreover, we present a novel Maximum Discrepancy Optimization (MaxDO) strategy to minimize the maximum discrepancy among multiple gradients. Solving by MaxDO with AGD, parallel training reduces the influence of the training conflict and suppresses the catastrophic forgetting of finished tasks. Extensive experiments validate the effectiveness of our approach on three image recognition datasets in task-incremental and class-incremental PCL. Our code is available at https://github.com/fanlyu/maxdo .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9cd7b0c0-cd6b-4d3b-9038-2f126d24223cCited by top-tier papers7
- Long-Tailed Learning as Multi-Objective OptimizationWeiqi Li, Fan Lyu, Fanhua Shang, Liang Wan et al.AAAI 2024 · 10 citations
- Rebalancing Multi-Label Class-Incremental LearningKaile Du, Yifan Zhou, Fan Lyu, Yuyang Li et al.AAAI 2025 · 6 citations
- FedAGC: Federated Continual Learning with Asymmetric Gradient CorrectionChengchao Zhang, Fanhua Shang, Hongyin Liu, Liang Wan et al.ICCV 2025 · 3 citations
- Beyond Myopic Alignment: Lookahead Optimization for Online Class-Incremental LearningSong Lai, Zhe Zhao, Fei Zhu, Ji Cheng et al.CVPR 2026
- Exposing Mixture and Annotating Confusion for Active Universal Test-Time AdaptationJiayao Tan, Fan Lyu, Chenggong Ni, Fuyuan Hu et al.ICLR 2026
Builds on9
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine et al.NeurIPS 2020 · 2,261 citations
- Conflict-Averse Gradient Descent for Multi-task learningBo Liu, Xingchao Liu, Xiaojie Jin, Peter Stone et al.NeurIPS 2021 · 686 citations
- Just Pick a Sign: Optimizing Deep Multitask Models with Gradient Sign DropoutZhao Chen, Jiquan Ngiam, Yanping Huang, Thang Luong et al.NeurIPS 2020 · 313 citations
- Federated Continual Learning with Weighted Inter-client TransferJaehong Yoon, Wonyong Jeong, Giwoong Lee, Eunho Yang et al.ICML 2021 · 303 citations
- Gradient Vaccine: Investigating and Improving Multi-task Optimization in Massively Multilingual ModelsZirui Wang, Yulia Tsvetkov, Orhan Firat, Yuan CaoICLR 2021 · 241 citations
Related papers
- Exploring The Forgetting in Adversarial Training: A Novel Method for Enhancing RobustnessXianglu Wang, Hu DingICLR 2025
- Data Augmented Flatness-aware Gradient Projection for Continual LearningEnneng Yang, Li Shen, Zhenyi Wang, Shiwei Liu et al.ICCV 2023 · 28 citations
- Class Gradient Projection For Continual LearningCheng Chen, Ji Zhang, Jingkuan Song, Lianli GaoACM MM 2022 · 14 citations
- Mitigating Catastrophic Forgetting in Online Continual Learning by Modeling Previous Task Interrelations via Pareto OptimizationYichen Wu, Hong Wang, Peilin Zhao, Yefeng Zheng et al.ICML 2024 · 23 citations
- Continual Learning by Using Information of Each Class HolisticallyWenpeng Hu, Qi Qin, Mengyu Wang, Jinwen Ma et al.AAAI 2021 · 64 citations
