No More Tuning: Prioritized Multi-Task Learning with Lagrangian Differential Multiplier Methods
Zhengxing Cheng, Yuheng Huang, Zhixuan Zhang, Dan Ou, Qingwen Liu
Abstract
Given the ubiquity of multi-task in practical systems, Multi-Task Learning (MTL) has found widespread application across diverse domains. In real-world scenarios, these tasks often have different priorities. For instance, In web search, relevance is often prioritized over other metrics, such as click-through rates or user engagement. Existing frameworks pay insufficient attention to the prioritization among different tasks, which typically adjust task-specific loss function weights to differentiate task priorities. However, this approach encounters challenges as the number of tasks grows, leading to exponential increases in hyper-parameter tuning complexity. Furthermore, the simultaneous optimization of multiple objectives can negatively impact the performance of high-priority tasks due to interference from lower-priority tasks. In this paper, we introduce a novel multi-task learning framework employing Lagrangian Differential Multiplier Methods for step-wise multi-task optimization. It is designed to boost the performance of high-priority tasks without interference from other tasks. Its primary advantage lies in its ability to automatically optimize multiple objectives without requiring balancing hyper-parameters for different tasks, thereby eliminating the need for manual tuning. Additionally, we provide theoretical analysis demonstrating that our method ensures optimization guarantees, enhancing the reliability of the process. We demonstrate its effectiveness through experiments on multiple public datasets and its application in Taobao search, a large-scale industrial search ranking system, resulting in significant improvements across various business metrics.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on7
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine et al.NeurIPS 2020 · 2,261 citations
- Responsive Safety in Reinforcement Learning by PID Lagrangian MethodsAdam Stooke, Joshua Achiam, Pieter AbbeelICML 2020 · 403 citations
- DSelect-k: Differentiable Selection in the Mixture of Experts with Applications to Multi-Task LearningHussein Hazimeh, Zhe Zhao, Aakanksha Chowdhery, Maheswaran Sathiamoorthy et al.NeurIPS 2021 · 216 citations
- FAMO: Fast Adaptive Multitask OptimizationBo Liu, Yihao Feng, Peter Stone, Qiang LiuNeurIPS 2023 · 127 citations
- RotoGrad: Gradient Homogenization in Multitask LearningAdrián Javaloy, Isabel ValeraICLR 2022 · 114 citations
Related papers
- Learning with Privileged TasksYuru Song, Zan Lou, Shan You, Erkun Yang et al.ICCV 2021 · 3 citations
- Aligned Multi Objective OptimizationYonathan Efroni, Ben Kretzu, Daniel Jiang, Jalaj Bhandari et al.ICML 2025
- Improving Gradient Trade-offs between Tasks in Multi-task Text ClassificationHeyan Chai, Jinhao Cui, Ye Wang, Min Zhang et al.ACL 2023 · 11 citations
- Quantifying Task Priority for Multi-Task OptimizationWooseong Jeong, Kuk-Jin YoonCVPR 2024
- Fair Resource Allocation in Multi-Task LearningHao Ban, Kaiyi JiICML 2024 · 41 citations
