MoDE: A Mixture-of-Experts Model with Mutual Distillation among the Experts
Zhitian Xie, Yinger Zhang, Chenyi Zhuang, Qitao Shi, Zhining Liu, Jinjie Gu, Guannan Zhang
Abstract
The application of mixture-of-experts (MoE) is gaining popularity due to its ability to improve model's performance. In an MoE structure, the gate layer plays a significant role in distinguishing and routing input features to different experts. This enables each expert to specialize in processing their corresponding sub-tasks. However, the gate's routing mechanism also gives rise to "narrow vision": the individual MoE's expert fails to use more samples in learning the allocated subtask, which in turn limits the MoE to further improve its generalization ability. To effectively address this, we propose a method called Mixture-of-Distilled-Expert (MoDE), which applies moderate mutual distillation among experts to enable each expert to pick up more features learned by other experts and gain more accurate perceptions on their allocated sub-tasks. We conduct plenty experiments including tabular, NLP and CV datasets, which shows MoDE's effectiveness, universality and robustness. Furthermore, we develop a parallel study through innovatively constructing "expert probing", to experimentally prove why MoDE works: moderate distilling knowledge from other experts can improve each individual expert's test performances on their assigned tasks, leading to MoE's overall performance improvement.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c64d488e-f59b-4183-862b-271dad945e0fCited by top-tier papers5
- Discovering Important Experts for Mixture-of-Experts Models Pruning Through a Theoretical PerspectiveWeizhong Huang, Yuxin Zhang, Xiawu Zheng, Fei Chao et al.NeurIPS 2025 · 12 citations
- MoE-GS: Mixture of Experts for Dynamic Gaussian SplattingIn-Hwan Jin, Hyeongju Mun, Joonsoo Kim, Kugjin Yun et al.ICLR 2026 · 2 citations
- Balanced Knowledge Distillation for Large Language Models with Mix-of-ExpertsJiajun Liu, Yao He, Wenjun Ke, Peng Wang et al.AAAI 2026
- Optimizing LoRA Allocation of MoE with the Alignment of Topic CorrelationHengyuan Xu, Wenjun Ke, Yao He, Jiajun Liu et al.AAAI 2026
- Combination-of-Experts with Knowledge Sharing for Cross-Task Vehicle Routing ProblemsZikang Yu, Jinbiao Chen, Jiahai WangICLR 2026
Builds on8
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen et al.ICLR 2021 · 1,954 citations
- Scaling Vision with Sparse Mixture of ExpertsCarlos Riquelme, Joan Puigcerver, Basil Mustafa, Maxim Neumann et al.NeurIPS 2021 · 1,213 citations
- BASE Layers: Simplifying Training of Large, Sparse ModelsMike Lewis, Shruti Bhosale, Tim Dettmers, Naman Goyal et al.ICML 2021 · 382 citations
- Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep LearningZeyuan Allen-Zhu, Yuanzhi LiICLR 2023 · 151 citations
Related papers
- HyperMoE: Towards Better Mixture of Experts via Transferring Among ExpertsHao Zhao, Zihan Qiu, Huijia Wu, Zili Wang et al.ACL 2024
- Theory on Mixture-of-Experts in Continual LearningHongbo Li, Sen Lin, Lingjie Duan, Yingbin Liang et al.ICLR 2025
- Dynamic Expert Specialization: Towards Catastrophic Forgetting-Free Multi-Domain MoE AdaptationJunzhuo Li, Bo Wang, Xiuze Zhou, Xuming HuEMNLP 2025 · 5 citations
- MoEC: Mixture of Expert ClustersYuan Xie, Shaohan Huang, Tianyu Chen, Furu WeiAAAI 2023 · 27 citations
- Efficient Deweahter Mixture-of-Experts with Uncertainty-Aware Feature-Wise Linear ModulationRongyu Zhang, Yulin Luo, Jiaming Liu, Huanrui Yang et al.AAAI 2024 · 30 citations
