DMTG: One-Shot Differentiable Multi-Task Grouping
Yuan Gao, Shuguo Jiang, Moran Li, Jin-Gang Yu, Gui-Song Xia
Abstract
We aim to address Multi-Task Learning (MTL) with a large number of tasks by Multi-Task Grouping (MTG). Given N tasks, we propose to simultaneously identify the best task groups from 2^N candidates and train the model weights simultaneously in one-shot, with the high-order task-affinity fully exploited. This is distinct from the pioneering methods which sequentially identify the groups and train the model weights, where the group identification often relies on heuristics. As a result, our method not only improves the training efficiency, but also mitigates the objective bias introduced by the sequential procedures that potentially lead to a suboptimal solution. Specifically, we formulate MTG as a fully differentiable pruning problem on an adaptive network architecture determined by an underlying Categorical distribution. To categorize N tasks into K groups (represented by K encoder branches), we initially set up KN task heads, where each branch connects to all N task heads to exploit the high-order task-affinity. Then, we gradually prune the KN heads down to N by learning a relaxed differentiable Categorical distribution, ensuring that each task is exclusively and uniquely categorized into only one branch. Extensive experiments on CelebA and Taskonomy datasets with detailed ablations show the promising performance and efficiency of our method. The codes are available at https://github.com/ethanygao/DMTG.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9be2d5ce-1088-44dc-9a45-e648b377ca6bCited by top-tier papers2
- MTRL-CG: Multi-Task Reinforcement Learning Method with Spectral Clustering-Based Task GroupingWenjia Meng, Teng Zhang, Haoliang Sun, Yilong YinAAAI 2026
- Few for Many: Tchebycheff Set Scalarization for Many-Objective OptimizationXi Lin, Yilu Liu, Xiaoyuan Zhang, Fei Liu et al.ICLR 2025
Builds on21
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Digging Into Self-Supervised Monocular Depth EstimationClément Godard, Oisin Mac Aodha, Michael Firman, Gabriel J. BrostowICCV 2019 · 2,416 citations
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine et al.NeurIPS 2020 · 2,261 citations
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 2,196 citations
- Conflict-Averse Gradient Descent for Multi-task learningBo Liu, Xingchao Liu, Xiaojie Jin, Peter Stone et al.NeurIPS 2021 · 686 citations
Related papers
- Efficiently Identifying Task Groupings for Multi-Task LearningChris Fifty, Ehsan Amid, Zhe Zhao, Tianhe Yu et al.NeurIPS 2021 · 352 citations
- Efficient and Effective Multi-task Grouping via Meta Learning on Task CombinationsXiaozhuang Song, Shun Zheng, Wei Cao, James J. Q. Yu et al.NeurIPS 2022 · 50 citations
- Learning to Branch for Multi-Task LearningPengsheng Guo, Chen-Yu Lee, Daniel UlbrichtICML 2020 · 208 citations
- Selective Task Group Updates for Multi-Task OptimizationWooseong Jeong, Kuk-Jin YoonICLR 2025
- Adaptive Activation Network and Functional Regularization for Efficient and Flexible Deep Multi-Task LearningYingru Liu, Xuewen Yang, Dongliang Xie, Xin Wang et al.AAAI 2020 · 10 citations
