Towards Consistent Multi-Task Learning: Unlocking the Potential of Task-Specific Parameters
Xiaohan Qin, Xiaoxing Wang, Junchi Yan
Abstract
Multi-task learning (MTL) has gained widespread application for its ability to transfer knowledge across tasks, improving resource efficiency and generalization. However, gradient conflicts from different tasks remain a major challenge in MTL. Previous gradient-based and lossbased methods primarily focus on gradient optimization in shared parameters, often overlooking the potential of taskspecific parameters. This work points out that task-specific parameters not only capture task-specific information but also influence the gradients propagated to shared parameters, which in turn affects gradient conflicts. Motivated by this insight, we propose ConsMTL, which models MTL as a bi-level optimization problem: in the upper-level optimization, we perform gradient aggregation on shared parameters to find a joint update vector that minimizes gradient conflicts; in the lower-level optimization, we introduce an additional loss for task-specific parameters guiding the k gradients of shared parameters to gradually converge towards the joint update vector. Our design enables the optimization of both shared and task-specific parameters to consistently mitigate gradient conflicts. Extensive experiments show that ConsMTL achieves state-of-the-art performance across various benchmarks with task numbers ranging from 2 to 40, demonstrating its superior performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b78a4b0d-9683-4f9e-852c-07d3ba82418aCited by top-tier papers5
- NTKMTL: Mitigating Task Imbalance in Multi-Task Learning from Neural Tangent Kernel PerspectiveXiaohan Qin, Xiaoxing Wang, Ning Liao, Junchi YanNeurIPS 2025 · 3 citations
- Stay in your Lane: Role Specific Queries with Overlap Suppression Loss for Dense Video CaptioningSeungHyup Baek, Jimin Lee, Hyeongkeun Lee, Jae Won ChoCVPR 2026 · 1 citation
- The Double Dilemma in Multi-Task Radiology Report Generation: A Gradient Dynamics Analysis and SolutionErjian Zhang, Yatong Hao, Liejun Wang, Zhiqing GuoICML 2026
- From Gradient Volume to Shapley Fairness: Towards Fair Multi-Task LearningXiao Wang, Yuying Han, Dazi Li, Fei Zhang et al.ICLR 2026
- Frequency Switching Mechanism for Parameter-Efficient Multi-Task LearningShih-Wen Liu, Yen-Chang Chen, Wei-Ta Chu, Fu-En Yang et al.CVPR 2026
Builds on14
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine et al.NeurIPS 2020 · 2,261 citations
- Conflict-Averse Gradient Descent for Multi-task learningBo Liu, Xingchao Liu, Xiaojie Jin, Peter Stone et al.NeurIPS 2021 · 686 citations
- Which Tasks Should Be Learned Together in Multi-task Learning?Trevor Standley, Amir Zamir, Dawn Chen, Leonidas J. Guibas et al.ICML 2020 · 651 citations
- Just Pick a Sign: Optimizing Deep Multitask Models with Gradient Sign DropoutZhao Chen, Jiquan Ngiam, Yanping Huang, Thang Luong et al.NeurIPS 2020 · 313 citations
- Multi-Task Learning as a Bargaining GameAviv Navon, Aviv Shamsian, Idan Achituve, Haggai Maron et al.ICML 2022 · 243 citations
Related papers
- Learning Conflict-Noticed Architecture for Multi-Task LearningZhixiong Yue, Yu Zhang, Jie LiangAAAI 2023 · 9 citations
- Adaptive Data and Task Joint Scheduling for Multi-Task LearningZeyu Liu, Heyan Chai, Chaoyang Li, Lingzhi Wang et al.ICDE 2025
- SAMO: A Lightweight Sharpness-Aware Approach for Multi-Task Optimization with Joint Global-Local PerturbationHao Ban, Gokul Ram Subramani, Kaiyi JiICCV 2025 · 3 citations
- Quantifying Task Priority for Multi-Task OptimizationWooseong Jeong, Kuk-Jin YoonCVPR 2024
- Improving Gradient Trade-offs between Tasks in Multi-task Text ClassificationHeyan Chai, Jinhao Cui, Ye Wang, Min Zhang et al.ACL 2023 · 11 citations
