Efficiently Identifying Task Groupings for Multi-Task Learning
Chris Fifty, Ehsan Amid, Zhe Zhao, Tianhe Yu, Rohan Anil, Chelsea Finn
摘要
Multi-task learning can leverage information learned by one task to benefit the training of other tasks. Despite this capacity, naïvely training all tasks together in one model often degrades performance, and exhaustively searching through combinations of task groupings can be prohibitively expensive. As a result, efficiently identifying the tasks that would benefit from training together remains a challenging design question without a clear solution. In this paper, we suggest an approach to select which tasks should train together in multi-task learning models. Our method determines task groupings in a single run by training all tasks together and quantifying the effect to which one task's gradient would affect another task's loss. On the large-scale Taskonomy computer vision dataset, we find this method can decrease test loss by 10.0% compared to simply training all tasks together while operating 11.6 times faster than a state-of-the-art task grouping method. Related Work Task Groupings. Prevailing wisdom suggests tasks which are similar or share a similar underlying structure may benefit from training together in a multi-task system [9, 8, 4] . Early work in this domain pertaining to the convex setting assume all tasks share a common latent feature representation, and find that model performance can be significantly improved by clustering tasks based on the basis vectors they share in this latent space [26, 30] . However, early convex methods to determine task groupings often make prohibitive assumptions that do not scale to deep neural networks. Deciding which tasks should train together in multi-task neural networks has traditionally been addressed with costly cross-validation techniques or high variance human intuition. An altogether
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper95
- TIES-Merging: Resolving Interference When Merging ModelsPrateek Yadav, Derek Tam, Leshem Choshen, Colin A. Raffel 等NeurIPS 2023 · 被引用 999 次
- ExT5: Towards Extreme Multi-Task Scaling for Transfer LearningVamsi Aribandi, Yi Tay, Tal Schuster, Jinfeng Rao 等ICLR 2022 · 被引用 237 次
- FAMO: Fast Adaptive Multitask OptimizationBo Liu, Yihao Feng, Peter Stone, Qiang LiuNeurIPS 2023 · 被引用 127 次
- RotoGrad: Gradient Homogenization in Multitask LearningAdrián Javaloy, Isabel ValeraICLR 2022 · 被引用 114 次
- Localizing Task Information for Improved Model Merging and CompressionKe Wang, Nikolaos Dimitriadis, Guillermo Ortiz-Jiménez, François Fleuret 等ICML 2024 · 被引用 107 次
它引用的顶会 Paper9
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine 等NeurIPS 2020 · 被引用 2,261 次
- Which Tasks Should Be Learned Together in Multi-task Learning?Trevor Standley, Amir Zamir, Dawn Chen, Leonidas J. Guibas 等ICML 2020 · 被引用 651 次
- Task2Vec: Task Embedding for Meta-LearningAlessandro Achille, Michael Lam, Rahul Tewari, Avinash Ravichandran 等ICCV 2019 · 被引用 359 次
- AdaShare: Learning What To Share For Efficient Deep Multi-Task LearningXimeng Sun, Rameswar Panda, Rogério Feris, Kate SaenkoNeurIPS 2020 · 被引用 337 次
- Just Pick a Sign: Optimizing Deep Multitask Models with Gradient Sign DropoutZhao Chen, Jiquan Ngiam, Yanping Huang, Thang Luong 等NeurIPS 2020 · 被引用 313 次
相关 Paper
- Efficient and Effective Multi-task Grouping via Meta Learning on Task CombinationsXiaozhuang Song, Shun Zheng, Wei Cao, James J. Q. Yu 等NeurIPS 2022 · 被引用 50 次
- DMTG: One-Shot Differentiable Multi-Task GroupingYuan Gao, Shuguo Jiang, Moran Li, Jin-Gang Yu 等ICML 2024 · 被引用 3 次
- Ensemble Prediction of Task Affinity for Efficient Multi-Task LearningAfiya Ayman, Ayan Mukhopadhyay, Aron LaszkaICLR 2026
- Selective Task Group Updates for Multi-Task OptimizationWooseong Jeong, Kuk-Jin YoonICLR 2025
- Adversarial Robustness in Multi-Task Learning: Promises and IllusionsSalah Ghamizi, Maxime Cordy, Mike Papadakis, Yves Le TraonAAAI 2022 · 被引用 25 次
