Measuring Asymmetric Gradient Discrepancy in Parallel Continual Learning
Fan Lyu, Qing Sun, Fanhua Shang, Liang Wan, Wei Feng
摘要
In Parallel Continual Learning (PCL), the parallel multiple tasks start and end training unpredictably, thus suffering from both training conflict and catastrophic forgetting issues. The two issues are raised because the gradients from parallel tasks differ in directions and magnitudes. Thus, in this paper, we formulate the PCL into a minimum distance optimization problem among gradients and propose an explicit Asymmetric Gradient Distance (AGD) to evaluate the gradient discrepancy in PCL. AGD considers both gradient magnitude ratios and directions, and has a tolerance when updating with a small gradient of inverse direction, which reduces the imbalanced influence of gradients on parallel task training. Moreover, we present a novel Maximum Discrepancy Optimization (MaxDO) strategy to minimize the maximum discrepancy among multiple gradients. Solving by MaxDO with AGD, parallel training reduces the influence of the training conflict and suppresses the catastrophic forgetting of finished tasks. Extensive experiments validate the effectiveness of our approach on three image recognition datasets in task-incremental and class-incremental PCL. Our code is available at https://github.com/fanlyu/maxdo .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Long-Tailed Learning as Multi-Objective OptimizationWeiqi Li, Fan Lyu, Fanhua Shang, Liang Wan 等AAAI 2024 · 被引用 10 次
- Rebalancing Multi-Label Class-Incremental LearningKaile Du, Yifan Zhou, Fan Lyu, Yuyang Li 等AAAI 2025 · 被引用 6 次
- FedAGC: Federated Continual Learning with Asymmetric Gradient CorrectionChengchao Zhang, Fanhua Shang, Hongyin Liu, Liang Wan 等ICCV 2025 · 被引用 3 次
- Beyond Myopic Alignment: Lookahead Optimization for Online Class-Incremental LearningSong Lai, Zhe Zhao, Fei Zhu, Ji Cheng 等CVPR 2026
- Exposing Mixture and Annotating Confusion for Active Universal Test-Time AdaptationJiayao Tan, Fan Lyu, Chenggong Ni, Fuyuan Hu 等ICLR 2026
它引用的顶会 Paper9
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine 等NeurIPS 2020 · 被引用 2,261 次
- Conflict-Averse Gradient Descent for Multi-task learningBo Liu, Xingchao Liu, Xiaojie Jin, Peter Stone 等NeurIPS 2021 · 被引用 686 次
- Just Pick a Sign: Optimizing Deep Multitask Models with Gradient Sign DropoutZhao Chen, Jiquan Ngiam, Yanping Huang, Thang Luong 等NeurIPS 2020 · 被引用 313 次
- Federated Continual Learning with Weighted Inter-client TransferJaehong Yoon, Wonyong Jeong, Giwoong Lee, Eunho Yang 等ICML 2021 · 被引用 303 次
- Gradient Vaccine: Investigating and Improving Multi-task Optimization in Massively Multilingual ModelsZirui Wang, Yulia Tsvetkov, Orhan Firat, Yuan CaoICLR 2021 · 被引用 241 次
相关 Paper
- Exploring The Forgetting in Adversarial Training: A Novel Method for Enhancing RobustnessXianglu Wang, Hu DingICLR 2025
- Data Augmented Flatness-aware Gradient Projection for Continual LearningEnneng Yang, Li Shen, Zhenyi Wang, Shiwei Liu 等ICCV 2023 · 被引用 28 次
- Class Gradient Projection For Continual LearningCheng Chen, Ji Zhang, Jingkuan Song, Lianli GaoACM MM 2022 · 被引用 14 次
- Mitigating Catastrophic Forgetting in Online Continual Learning by Modeling Previous Task Interrelations via Pareto OptimizationYichen Wu, Hong Wang, Peilin Zhao, Yefeng Zheng 等ICML 2024 · 被引用 23 次
- Continual Learning by Using Information of Each Class HolisticallyWenpeng Hu, Qi Qin, Mengyu Wang, Jinwen Ma 等AAAI 2021 · 被引用 64 次
