RotoGrad: Gradient Homogenization in Multitask Learning
Adrián Javaloy, Isabel Valera
摘要
Multitask learning is being increasingly adopted in applications domains like computer vision and reinforcement learning. However, optimally exploiting its advantages remains a major challenge due to the effect of negative transfer. Previous works have tracked down this issue to the disparities in gradient magnitudes and directions across tasks when optimizing the shared network parameters. While recent work has acknowledged that negative transfer is a two-fold problem, existing approaches fall short. These methods only focus on either homogenizing the gradient magnitude across tasks; or greedily change the gradient directions, overlooking future conflicts. In this work, we introduce RotoGrad, an algorithm that tackles negative transfer as a whole: it jointly homogenizes gradient magnitudes and directions, while ensuring training convergence. We show that RotoGrad outperforms competing methods in complex problems, including multi-label classification in CelebA and computer vision tasks in the NYUv2 dataset. A Pytorch implementation can be found in https://github.com/adrianjav/rotograd .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper54
- Representation Surgery for Multi-Task Model MergingEnneng Yang, Li Shen, Zhenyi Wang, Guibing Guo 等ICML 2024 · 被引用 96 次
- In Defense of the Unitary Scalarization for Deep Multi-Task LearningVitaly Kurin, Alessandro De Palma, Ilya Kostrikov, Shimon Whiteson 等NeurIPS 2022 · 被引用 96 次
- On the Convergence of Stochastic Multi-Objective Gradient Manipulation and BeyondShiji Zhou, Wenpeng Zhang, Jiyan Jiang, Wenliang Zhong 等NeurIPS 2022 · 被引用 66 次
- ForkMerge: Mitigating Negative Transfer in Auxiliary-Task LearningJunguang Jiang, Baixu Chen, Junwei Pan, Ximei Wang 等NeurIPS 2023 · 被引用 55 次
- Direction-oriented Multi-objective Learning: Simple and Provable Stochastic AlgorithmsPeiyao Xiao, Hao Ban, Kaiyi JiNeurIPS 2023 · 被引用 46 次
它引用的顶会 Paper14
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine 等NeurIPS 2020 · 被引用 2,261 次
- On the Variance of the Adaptive Learning Rate and BeyondLiyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen 等ICLR 2020 · 被引用 2,210 次
- Conflict-Averse Gradient Descent for Multi-task learningBo Liu, Xingchao Liu, Xiaojie Jin, Peter Stone 等NeurIPS 2021 · 被引用 686 次
- Which Tasks Should Be Learned Together in Multi-task Learning?Trevor Standley, Amir Zamir, Dawn Chen, Leonidas J. Guibas 等ICML 2020 · 被引用 651 次
- What is Local Optimality in Nonconvex-Nonconcave Minimax Optimization?Chi Jin, Praneeth Netrapalli, Michael I. JordanICML 2020 · 被引用 381 次
相关 Paper
- Towards Task-Conflicts Momentum-Calibrated Approach for Multi-task LearningHeyan Chai, Zeyu Liu, Yongxin Tong, Ziyi Yao 等ICDE 2024 · 被引用 5 次
- Fair Resource Allocation in Multi-Task LearningHao Ban, Kaiyi JiICML 2024 · 被引用 41 次
- Learning Conflict-Noticed Architecture for Multi-Task LearningZhixiong Yue, Yu Zhang, Jie LiangAAAI 2023 · 被引用 9 次
- TaskForce: Cooperative Multi-agent Reinforcement Learning for Multi-task OptimizationWonhyeok Choi, Kyumin Hwang, Jihun Park, Kyoungmin Lee 等CVPR 2026
- Towards Consistent Multi-Task Learning: Unlocking the Potential of Task-Specific ParametersXiaohan Qin, Xiaoxing Wang, Junchi YanCVPR 2025
