Robust Multi-Task Learning with Excess Risks
Yifei He, Shiji Zhou, Guojun Zhang, Hyokun Yun, Yi Xu, Belinda Zeng, Trishul Chilimbi, Han Zhao
摘要
Multi-task learning (MTL) considers learning a joint model for multiple tasks by optimizing a convex combination of all task losses. To solve the optimization problem, existing methods use an adaptive weight updating scheme, where task weights are dynamically adjusted based on their respective losses to prioritize difficult tasks. However, these algorithms face a great challenge whenever label noise is present, in which case excessive weights tend to be assigned to noisy tasks that have relatively large Bayes optimal errors, thereby overshadowing other tasks and causing performance to drop across the board. To overcome this limitation, we propose Multi-Task Learning with Excess Risks (ExcessMTL), an excess risk-based task balancing method that updates the task weights by their distances to convergence instead. Intuitively, ExcessMTL assigns higher weights to worse-trained tasks that are further from convergence. To estimate the excess risks, we develop an efficient and accurate method with Taylor approximation. Theoretically, we show that our proposed algorithm achieves convergence guarantees and Pareto stationarity. Empirically, we evaluate our algorithm on various MTL benchmarks and demonstrate its superior performance over existing methods in the presence of label noise. Our code is available at https://github.com/yifei-he/ExcessMTL .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Efficient Utility-Preserving Machine Unlearning with Implicit Gradient SurgeryShiji Zhou, Tianbai Yu, Zhi Zhang, Heng Chang 等NeurIPS 2025 · 被引用 6 次
- On the Plasticity and Stability for Post-Training Large Language ModelsWenwen Qiang, Ziyin Gu, Jiahuan Zhou, Jie Hu 等ICML 2026 · 被引用 3 次
- CORE-MTL: Rethinking Gradient Balancing via Causal Orthogonal RepresentationsChengfeng Wu, Tao Zou, Yanru Wu, Jingge WangICML 2026
- AutoMixAlign: Adaptive Data Mixing for Multi-Task Preference Optimization in LLMsNicholas E. Corrado, Julian Katz-Samuels, Adithya M. Devraj, Hyokun Yun 等ACL 2025
- LoRA-DA: Data-Aware Initialization for Low-Rank Adaptation via Asymptotic AnalysisQingyue Zhang, Chang Chu, Tianren Peng, Qi Li 等ICML 2026
它引用的顶会 Paper15
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine 等NeurIPS 2020 · 被引用 2,261 次
- Distributionally Robust Neural NetworksShiori Sagawa, Pang Wei Koh, Tatsunori B. Hashimoto, Percy LiangICLR 2020 · 被引用 1,578 次
- Conflict-Averse Gradient Descent for Multi-task learningBo Liu, Xingchao Liu, Xiaojie Jin, Peter Stone 等NeurIPS 2021 · 被引用 686 次
- Just Train Twice: Improving Group Robustness without Training Group InformationEvan Zheran Liu, Behzad Haghgoo, Annie S. Chen, Aditi Raghunathan 等ICML 2021 · 被引用 683 次
- NLNL: Negative Learning for Noisy LabelsYoungdong Kim, Junho Yim, Juseung Yun, Junmo KimICCV 2019 · 被引用 338 次
相关 Paper
- Learning Multiple Pixelwise Tasks Based on Loss Scale BalancingJae-Han Lee, Chul Lee, Chang-Su KimICCV 2021 · 被引用 13 次
- Multi-Task Learning with User Preferences: Gradient Descent with Controlled Ascent in Pareto OptimizationDebabrata Mahapatra, Vaibhav RajanICML 2020 · 被引用 182 次
- Multi-Task Learning as a Bargaining GameAviv Navon, Aviv Shamsian, Idan Achituve, Haggai Maron 等ICML 2022 · 被引用 243 次
- Multi-Task Representation Alignment on Language Understanding: A Mutual Information PerspectiveDou Hu, Lingwei Wei, Hongjiang Xiao, Songlin Hu 等ACL 2026
- Revisiting Fairness in Multitask Learning: A Performance-Driven Approach for Variance ReductionXiaohan Qin, Xiaoxing Wang, Junchi YanCVPR 2025
