Robust Multi-Task Learning with Excess Risks
Yifei He, Shiji Zhou, Guojun Zhang, Hyokun Yun, Yi Xu, Belinda Zeng, Trishul Chilimbi, Han Zhao
Abstract
Multi-task learning (MTL) considers learning a joint model for multiple tasks by optimizing a convex combination of all task losses. To solve the optimization problem, existing methods use an adaptive weight updating scheme, where task weights are dynamically adjusted based on their respective losses to prioritize difficult tasks. However, these algorithms face a great challenge whenever label noise is present, in which case excessive weights tend to be assigned to noisy tasks that have relatively large Bayes optimal errors, thereby overshadowing other tasks and causing performance to drop across the board. To overcome this limitation, we propose Multi-Task Learning with Excess Risks (ExcessMTL), an excess risk-based task balancing method that updates the task weights by their distances to convergence instead. Intuitively, ExcessMTL assigns higher weights to worse-trained tasks that are further from convergence. To estimate the excess risks, we develop an efficient and accurate method with Taylor approximation. Theoretically, we show that our proposed algorithm achieves convergence guarantees and Pareto stationarity. Empirically, we evaluate our algorithm on various MTL benchmarks and demonstrate its superior performance over existing methods in the presence of label noise. Our code is available at https://github.com/yifei-he/ExcessMTL .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4d9edd24-7324-40d3-a5e9-6502a0693132Cited by top-tier papers8
- Efficient Utility-Preserving Machine Unlearning with Implicit Gradient SurgeryShiji Zhou, Tianbai Yu, Zhi Zhang, Heng Chang et al.NeurIPS 2025 · 6 citations
- On the Plasticity and Stability for Post-Training Large Language ModelsWenwen Qiang, Ziyin Gu, Jiahuan Zhou, Jie Hu et al.ICML 2026 · 3 citations
- CORE-MTL: Rethinking Gradient Balancing via Causal Orthogonal RepresentationsChengfeng Wu, Tao Zou, Yanru Wu, Jingge WangICML 2026
- AutoMixAlign: Adaptive Data Mixing for Multi-Task Preference Optimization in LLMsNicholas E. Corrado, Julian Katz-Samuels, Adithya M. Devraj, Hyokun Yun et al.ACL 2025
- LoRA-DA: Data-Aware Initialization for Low-Rank Adaptation via Asymptotic AnalysisQingyue Zhang, Chang Chu, Tianren Peng, Qi Li et al.ICML 2026
Builds on15
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine et al.NeurIPS 2020 · 2,261 citations
- Distributionally Robust Neural NetworksShiori Sagawa, Pang Wei Koh, Tatsunori B. Hashimoto, Percy LiangICLR 2020 · 1,578 citations
- Conflict-Averse Gradient Descent for Multi-task learningBo Liu, Xingchao Liu, Xiaojie Jin, Peter Stone et al.NeurIPS 2021 · 686 citations
- Just Train Twice: Improving Group Robustness without Training Group InformationEvan Zheran Liu, Behzad Haghgoo, Annie S. Chen, Aditi Raghunathan et al.ICML 2021 · 683 citations
- NLNL: Negative Learning for Noisy LabelsYoungdong Kim, Junho Yim, Juseung Yun, Junmo KimICCV 2019 · 338 citations
Related papers
- Learning Multiple Pixelwise Tasks Based on Loss Scale BalancingJae-Han Lee, Chul Lee, Chang-Su KimICCV 2021 · 13 citations
- Multi-Task Learning with User Preferences: Gradient Descent with Controlled Ascent in Pareto OptimizationDebabrata Mahapatra, Vaibhav RajanICML 2020 · 182 citations
- Multi-Task Learning as a Bargaining GameAviv Navon, Aviv Shamsian, Idan Achituve, Haggai Maron et al.ICML 2022 · 243 citations
- Multi-Task Representation Alignment on Language Understanding: A Mutual Information PerspectiveDou Hu, Lingwei Wei, Hongjiang Xiao, Songlin Hu et al.ACL 2026
- Revisiting Fairness in Multitask Learning: A Performance-Driven Approach for Variance ReductionXiaohan Qin, Xiaoxing Wang, Junchi YanCVPR 2025
