Mod-Squad: Designing Mixtures of Experts As Modular Multi-Task Learners
Zitian Chen, Yikang Shen, Mingyu Ding, Zhenfang Chen, Hengshuang Zhao, Erik G. Learned-Miller, Chuang Gan
Abstract
Optimization in multi-task learning (MTL) is more challenging than single-task learning (STL), as the gradient from different tasks can be contradictory. When tasks are related, it can be beneficial to share some parameters among them (cooperation). However, some tasks require additional parameters with expertise in a specific type of data or discrimination (specialization). To address the MTL challenge, we propose Mod-Squad, a new model that is Modularized into groups of experts (a 'Squad'). This structure allows us to formalize cooperation and specialization as the process of matching experts and tasks. We optimize this matching process during the training of a single model. Specifically, we incorporate mixture of experts (MoE) layers into a transformer model, with a new loss that incorporates the mutual dependence between tasks and experts. As a result, only a small set of experts are activated for each task. This prevents the sharing of the entire backbone model between all tasks, which strengthens the model, especially when the training set size and the number of tasks scale up. More interestingly, for each task, we can extract the small set of experts as a standalone model that maintains the same performance as the large model. Extensive experiments on the Taskonomy dataset with 13 vision tasks and the PASCAL-Context dataset with 5 vision tasks show the superiority of our approach. The project page can be accessed at https://vis-www.cs.umass.edu/mod-squad .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 15aa59e9-3a62-4dff-9fc5-5be479bfddf0Cited by top-tier papers52
- Flex-MoE: Modeling Arbitrary Modality Combination via the Flexible Mixture-of-ExpertsSukwon Yun, Inyoung Choi, Jie Peng, Yangfan Wu et al.NeurIPS 2024 · 98 citations
- Boosting Continual Learning of Vision-Language Models via Mixture-of-Experts AdaptersJiazuo Yu, Yunzhi Zhuge, Lu Zhang, Ping Hu et al.CVPR 2024 · 80 citations
- Mixtures of Experts Unlock Parameter Scaling for Deep RLJohan S. Obando-Ceron, Ghada Sokar, Timon Willi, Clare Lyle et al.ICML 2024 · 74 citations
- DINO-Foresight: Looking into the Future with DINOEfstathios Karypidis, Ioannis Kakogeorgiou, Spyridon Gidaris, Nikos KomodakisNeurIPS 2025 · 52 citations
- Multi-Task Dense Prediction via Mixture of Low-Rank ExpertsYuqi Yang, Peng-Tao Jiang, Qibin Hou, Hao Zhang et al.CVPR 2024 · 32 citations
Builds on15
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine et al.NeurIPS 2020 · 2,261 citations
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen et al.ICLR 2021 · 1,954 citations
- Once-for-All: Train One Network and Specialize it for Efficient DeploymentHan Cai, Chuang Gan, Tianzhe Wang, Zhekai Zhang et al.ICLR 2020 · 1,522 citations
- Scaling Vision with Sparse Mixture of ExpertsCarlos Riquelme, Joan Puigcerver, Basil Mustafa, Maxim Neumann et al.NeurIPS 2021 · 1,213 citations
Related papers
- AdaMV-MoE: Adaptive Multi-Task Vision Mixture-of-ExpertsTianlong Chen, Xuxi Chen, Xianzhi Du, Abdullah Rashwan et al.ICCV 2023 · 119 citations
- Multi-Task Dense Prediction Fine-Tuning with Mixture of Fine-Grained ExpertsYangyang Xu, Xi Ye, Duo SuACM MM 2025
- Synergistic Intra- and Cross-Layer Regularization Losses for MoE Expert SpecializationRizhen Hu, Yuan Cao, Boao Kong, Mou Sun et al.ICML 2026
- Merging Multi-Task Models via Weight-Ensembling Mixture of ExpertsAnke Tang, Li Shen, Yong Luo, Nan Yin et al.ICML 2024 · 96 citations
- Each Rank Could be an Expert: Single-Ranked Mixture of Experts LoRA for Multi-task LearningZiyu Zhao, Yixiao Zhou, Xin Yu, Zhi Zhang et al.KDD 2026 · 13 citations
