Towards Consistent Multi-Task Learning: Unlocking the Potential of Task-Specific Parameters
Xiaohan Qin, Xiaoxing Wang, Junchi Yan
摘要
Multi-task learning (MTL) has gained widespread application for its ability to transfer knowledge across tasks, improving resource efficiency and generalization. However, gradient conflicts from different tasks remain a major challenge in MTL. Previous gradient-based and lossbased methods primarily focus on gradient optimization in shared parameters, often overlooking the potential of taskspecific parameters. This work points out that task-specific parameters not only capture task-specific information but also influence the gradients propagated to shared parameters, which in turn affects gradient conflicts. Motivated by this insight, we propose ConsMTL, which models MTL as a bi-level optimization problem: in the upper-level optimization, we perform gradient aggregation on shared parameters to find a joint update vector that minimizes gradient conflicts; in the lower-level optimization, we introduce an additional loss for task-specific parameters guiding the k gradients of shared parameters to gradually converge towards the joint update vector. Our design enables the optimization of both shared and task-specific parameters to consistently mitigate gradient conflicts. Extensive experiments show that ConsMTL achieves state-of-the-art performance across various benchmarks with task numbers ranging from 2 to 40, demonstrating its superior performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- NTKMTL: Mitigating Task Imbalance in Multi-Task Learning from Neural Tangent Kernel PerspectiveXiaohan Qin, Xiaoxing Wang, Ning Liao, Junchi YanNeurIPS 2025 · 被引用 3 次
- Stay in your Lane: Role Specific Queries with Overlap Suppression Loss for Dense Video CaptioningSeungHyup Baek, Jimin Lee, Hyeongkeun Lee, Jae Won ChoCVPR 2026 · 被引用 1 次
- The Double Dilemma in Multi-Task Radiology Report Generation: A Gradient Dynamics Analysis and SolutionErjian Zhang, Yatong Hao, Liejun Wang, Zhiqing GuoICML 2026
- From Gradient Volume to Shapley Fairness: Towards Fair Multi-Task LearningXiao Wang, Yuying Han, Dazi Li, Fei Zhang 等ICLR 2026
- Frequency Switching Mechanism for Parameter-Efficient Multi-Task LearningShih-Wen Liu, Yen-Chang Chen, Wei-Ta Chu, Fu-En Yang 等CVPR 2026
它引用的顶会 Paper14
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine 等NeurIPS 2020 · 被引用 2,261 次
- Conflict-Averse Gradient Descent for Multi-task learningBo Liu, Xingchao Liu, Xiaojie Jin, Peter Stone 等NeurIPS 2021 · 被引用 686 次
- Which Tasks Should Be Learned Together in Multi-task Learning?Trevor Standley, Amir Zamir, Dawn Chen, Leonidas J. Guibas 等ICML 2020 · 被引用 651 次
- Just Pick a Sign: Optimizing Deep Multitask Models with Gradient Sign DropoutZhao Chen, Jiquan Ngiam, Yanping Huang, Thang Luong 等NeurIPS 2020 · 被引用 313 次
- Multi-Task Learning as a Bargaining GameAviv Navon, Aviv Shamsian, Idan Achituve, Haggai Maron 等ICML 2022 · 被引用 243 次
相关 Paper
- Learning Conflict-Noticed Architecture for Multi-Task LearningZhixiong Yue, Yu Zhang, Jie LiangAAAI 2023 · 被引用 9 次
- Adaptive Data and Task Joint Scheduling for Multi-Task LearningZeyu Liu, Heyan Chai, Chaoyang Li, Lingzhi Wang 等ICDE 2025
- SAMO: A Lightweight Sharpness-Aware Approach for Multi-Task Optimization with Joint Global-Local PerturbationHao Ban, Gokul Ram Subramani, Kaiyi JiICCV 2025 · 被引用 3 次
- Quantifying Task Priority for Multi-Task OptimizationWooseong Jeong, Kuk-Jin YoonCVPR 2024
- Improving Gradient Trade-offs between Tasks in Multi-task Text ClassificationHeyan Chai, Jinhao Cui, Ye Wang, Min Zhang 等ACL 2023 · 被引用 11 次
