AdaTask: A Task-Aware Adaptive Learning Rate Approach to Multi-Task Learning
Enneng Yang, Junwei Pan, Ximei Wang, Haibin Yu, Li Shen, Xihua Chen, Lei Xiao, Jie Jiang, Guibing Guo
Abstract
Multi-task learning (MTL) models have demonstrated impressive results in computer vision, natural language processing, and recommender systems. Even though many approaches have been proposed, how well these approaches balance different tasks on each parameter still remains unclear. In this paper, we propose to measure the task dominance degree of a parameter by the total updates of each task on this parameter. Specifically, we compute the total updates by the exponentially decaying Average of the squared Updates (AU) on a parameter from the corresponding task. Based on this novel metric, we observe that many parameters in existing MTL methods, especially those in the higher shared layers, are still dominated by one or several tasks. The dominance of AU is mainly due to the dominance of accumulative gradients from one or several tasks. Motivated by this, we propose a Task-wise Adaptive learning rate approach, AdaTask in short, to separate the accumulative gradients and hence the learning rate of each task for each parameter in adaptive learning rate approaches (e.g., AdaGrad, RMSProp, and Adam). Comprehensive experiments on computer vision and recommender system MTL datasets demonstrate that AdaTask significantly improves the performance of dominated tasks, resulting SOTA average task-wise performance. Analysis on both synthetic and real-world datasets shows AdaTask balance parameters in every shared layer well.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1706b1c9-7af0-4366-9cfd-d320d2eb739cCited by top-tier papers21
- AdaMerging: Adaptive Model Merging for Multi-Task LearningEnneng Yang, Zhenyi Wang, Li Shen, Shiwei Liu et al.ICLR 2024 · 230 citations
- Representation Surgery for Multi-Task Model MergingEnneng Yang, Li Shen, Zhenyi Wang, Guibing Guo et al.ICML 2024 · 96 citations
- Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task LearnersMichal Nauman, Marek Cygan, Carmelo Sferrazza, Aviral Kumar et al.NeurIPS 2025 · 26 citations
- D3: A Methodological Exploration of Domain Division, Modeling, and Balance in Multi-Domain RecommendationsPengyue Jia, Yichao Wang, Shanru Lin, Xiaopeng Li et al.AAAI 2024 · 13 citations
- PiKE: Adaptive Data Mixing for Large-Scale Multi-Task Learning Under Low Gradient ConflictsZeman Li, Yuan Deng, Peilin Zhong, Meisam Razaviyayn et al.NeurIPS 2025 · 8 citations
Builds on10
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine et al.NeurIPS 2020 · 2,261 citations
- AdaBelief Optimizer: Adapting Stepsizes by the Belief in Observed GradientsJuntang Zhuang, Tommy Tang, Yifan Ding, Sekhar Tatikonda et al.NeurIPS 2020 · 697 citations
- Conflict-Averse Gradient Descent for Multi-task learningBo Liu, Xingchao Liu, Xiaojie Jin, Peter Stone et al.NeurIPS 2021 · 686 citations
- AdaShare: Learning What To Share For Efficient Deep Multi-Task LearningXimeng Sun, Rameswar Panda, Rogério Feris, Kate SaenkoNeurIPS 2020 · 337 citations
- Just Pick a Sign: Optimizing Deep Multitask Models with Gradient Sign DropoutZhao Chen, Jiquan Ngiam, Yanping Huang, Thang Luong et al.NeurIPS 2020 · 313 citations
Related papers
- Learning Multiple Pixelwise Tasks Based on Loss Scale BalancingJae-Han Lee, Chul Lee, Chang-Su KimICCV 2021 · 13 citations
- MTAdam: Automatic Balancing of Multiple Training Loss TermsItzik Malkiel, Lior WolfEMNLP 2021 · 12 citations
- FAMO: Fast Adaptive Multitask OptimizationBo Liu, Yihao Feng, Peter Stone, Qiang LiuNeurIPS 2023 · 127 citations
- Learn to Merge: Meta-Learning for Adaptive Multi-Task Model MergingJun Chen, Qin Zhang, Weizhi Zhang, Xiao Luo et al.ICML 2026
- Towards Consistent Multi-Task Learning: Unlocking the Potential of Task-Specific ParametersXiaohan Qin, Xiaoxing Wang, Junchi YanCVPR 2025
