MoDE: A Mixture-of-Experts Model with Mutual Distillation among the Experts
Zhitian Xie, Yinger Zhang, Chenyi Zhuang, Qitao Shi, Zhining Liu, Jinjie Gu, Guannan Zhang
摘要
The application of mixture-of-experts (MoE) is gaining popularity due to its ability to improve model's performance. In an MoE structure, the gate layer plays a significant role in distinguishing and routing input features to different experts. This enables each expert to specialize in processing their corresponding sub-tasks. However, the gate's routing mechanism also gives rise to "narrow vision": the individual MoE's expert fails to use more samples in learning the allocated subtask, which in turn limits the MoE to further improve its generalization ability. To effectively address this, we propose a method called Mixture-of-Distilled-Expert (MoDE), which applies moderate mutual distillation among experts to enable each expert to pick up more features learned by other experts and gain more accurate perceptions on their allocated sub-tasks. We conduct plenty experiments including tabular, NLP and CV datasets, which shows MoDE's effectiveness, universality and robustness. Furthermore, we develop a parallel study through innovatively constructing "expert probing", to experimentally prove why MoDE works: moderate distilling knowledge from other experts can improve each individual expert's test performances on their assigned tasks, leading to MoE's overall performance improvement.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Discovering Important Experts for Mixture-of-Experts Models Pruning Through a Theoretical PerspectiveWeizhong Huang, Yuxin Zhang, Xiawu Zheng, Fei Chao 等NeurIPS 2025 · 被引用 12 次
- MoE-GS: Mixture of Experts for Dynamic Gaussian SplattingIn-Hwan Jin, Hyeongju Mun, Joonsoo Kim, Kugjin Yun 等ICLR 2026 · 被引用 2 次
- Balanced Knowledge Distillation for Large Language Models with Mix-of-ExpertsJiajun Liu, Yao He, Wenjun Ke, Peng Wang 等AAAI 2026
- Optimizing LoRA Allocation of MoE with the Alignment of Topic CorrelationHengyuan Xu, Wenjun Ke, Yao He, Jiajun Liu 等AAAI 2026
- Combination-of-Experts with Knowledge Sharing for Cross-Task Vehicle Routing ProblemsZikang Yu, Jinbiao Chen, Jiahai WangICLR 2026
它引用的顶会 Paper8
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen 等ICLR 2021 · 被引用 1,954 次
- Scaling Vision with Sparse Mixture of ExpertsCarlos Riquelme, Joan Puigcerver, Basil Mustafa, Maxim Neumann 等NeurIPS 2021 · 被引用 1,213 次
- BASE Layers: Simplifying Training of Large, Sparse ModelsMike Lewis, Shruti Bhosale, Tim Dettmers, Naman Goyal 等ICML 2021 · 被引用 382 次
- Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep LearningZeyuan Allen-Zhu, Yuanzhi LiICLR 2023 · 被引用 151 次
相关 Paper
- HyperMoE: Towards Better Mixture of Experts via Transferring Among ExpertsHao Zhao, Zihan Qiu, Huijia Wu, Zili Wang 等ACL 2024
- Theory on Mixture-of-Experts in Continual LearningHongbo Li, Sen Lin, Lingjie Duan, Yingbin Liang 等ICLR 2025
- Dynamic Expert Specialization: Towards Catastrophic Forgetting-Free Multi-Domain MoE AdaptationJunzhuo Li, Bo Wang, Xiuze Zhou, Xuming HuEMNLP 2025 · 被引用 5 次
- MoEC: Mixture of Expert ClustersYuan Xie, Shaohan Huang, Tianyu Chen, Furu WeiAAAI 2023 · 被引用 27 次
- Efficient Deweahter Mixture-of-Experts with Uncertainty-Aware Feature-Wise Linear ModulationRongyu Zhang, Yulin Luo, Jiaming Liu, Huanrui Yang 等AAAI 2024 · 被引用 30 次
