Bagging-Expert Network for Multi-Task Learning: A Depolarization Solution in Multi-Gate Mixture-of-Experts
Gong-Duo Zhang, Ruiqing Chen, Qian Zhao, Zhengwei Wu, Fengyu Han, Huan-Yi Su, Ziqi Liu, Lihong Gu, Lin Zhou
Abstract
Multi-task learning (MTL) is widely utilized across a variety of real-world applications, including recommendation systems. For instance, in the field of e-commerce, MTL is commonly employed to simultaneously model click, conversion, and user dwelling time. Among a various of MTL models, the Multi-gate Mixture-of-Experts (MMoE) has gained significant popularity. However, MMoE suffers from the polarization issue during training, where the weights of certain experts tend to converge towards 0. To address this issue, we propose a novel method called Bagging-Expert network (BEnet) for multi-task learning. BEnet effectively mitigates the problem of polarization and achieves excellent performance in multi-task learning. It incorporates a bagging layer and an attention mechanism to encourage experts focusing on diverse knowledge domains. Simultaneously, polarization is avoided as different experts execute respective duties and specialize in distinct domains. Experimental results on real-world datasets demonstrate that BEnet has strong robustness and outperforms other state-of-the-art (SOTA) MTL methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2bbdf68c-2646-4144-8070-d9dff1e2d3a7Builds on7
- GLaM: Efficient Scaling of Language Models with Mixture-of-ExpertsNan Du, Yanping Huang, Andrew M. Dai, Simon Tong et al.ICML 2022 · 1,173 citations
- Mixture-of-Experts with Expert Choice RoutingYanqi Zhou, Tao Lei, Hanxiao Liu, Nan Du et al.NeurIPS 2022 · 933 citations
- DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language ModelsDamai Dai, Chengqi Deng, Chenggang Zhao, R. X. Xu et al.ACL 2024 · 171 citations
- SummaReranker: A Multi-Task Mixture-of-Experts Re-ranking Framework for Abstractive SummarizationMathieu Ravaut, Shafiq R. Joty, Nancy F. ChenACL 2022 · 116 citations
- Octavius: Mitigating Task Interference in MLLMs via LoRA-MoEZeren Chen, Ziqin Wang, Zhen Wang, Huayang Liu et al.ICLR 2024 · 24 citations
Related papers
- Automatic Expert Selection for Multi-Scenario and Multi-Task SearchXinyu Zou, Zhi Hu, Yiming Zhao, Xuchu Ding et al.SIGIR 2022 · 43 citations
- M3oE: Multi-Domain Multi-Task Mixture-of Experts Recommendation FrameworkZijian Zhang, Shuchang Liu, Jiaao Yu, Qingpeng Cai et al.SIGIR 2024 · 27 citations
- Theory on Mixture-of-Experts in Continual LearningHongbo Li, Sen Lin, Lingjie Duan, Yingbin Liang et al.ICLR 2025
- DSelect-k: Differentiable Selection in the Mixture of Experts with Applications to Multi-Task LearningHussein Hazimeh, Zhe Zhao, Aakanksha Chowdhery, Maheswaran Sathiamoorthy et al.NeurIPS 2021 · 216 citations
- Multi-Task Dense Prediction Fine-Tuning with Mixture of Fine-Grained ExpertsYangyang Xu, Xi Ye, Duo SuACM MM 2025
