Teacher-Guided Routing for Sparse Vision Mixture-of-Experts
Masahiro Kada, Ryota Yoshihashi, Satoshi Ikehata, Rei Kawakami, Ikuro Sato
Abstract
Recent progress in deep learning has been driven by increasingly large-scale models, but the resulting computational cost has become a critical bottleneck.Sparse Mixture of Experts (MoE) offers an effective solution by activating only a small subset of expert networks for each input, achieving high scalability with limited computation.Although effective, sparse MoE training exhibits characteristic optimization difficulties. Because the router receives gradients only from the experts it selects in each forward pass, its learning signal is highly localized, with little information about the broader expert space.This limited gradient feedback can lead the router toward suboptimal configurations, for example collapsing to only a few experts when no auxiliary losses are used, and it has also been associated with fluctuating expert selections during training. These behaviors suggest that task-driven signals alone do not provide sufficient guidance for learning robust routing behavior in sparse MoE.To address this issue, we propose TGR-MoE: Teacher-Guided Routing for Sparse Vision Mixture-of-Experts, a simple yet effective method that stabilizes router learning using supervision derived from a pretrained dense teacher model.TGR-MoE constructs a teacher router from the teacher's intermediate representations and uses its routing outputs as pseudo-supervision for the student router, suppressing frequent routing fluctuations during training and enabling knowledge-guided expert selection from the early stages of training.Extensive experiments on ImageNet-1K and CIFAR-100 demonstrate that TGR consistently improves both accuracy and routing consistency, while maintaining stable training even under highly sparse configurations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0446fc92-8e53-4d37-84ae-8f1da933a327Builds on26
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh et al.ICCV 2019 · 5,843 citations
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 4,453 citations
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen et al.ICLR 2021 · 1,954 citations
- Scaling Vision with Sparse Mixture of ExpertsCarlos Riquelme, Joan Puigcerver, Basil Mustafa, Maxim Neumann et al.NeurIPS 2021 · 1,213 citations
Related papers
- A Provably Effective Method for Pruning Experts in Fine-tuned Sparse Mixture-of-ExpertsMohammed Nowaz Rabbani Chowdhury, Meng Wang, Kaoutar El Maghraoui, Naigang Wang et al.ICML 2024 · 18 citations
- Dense Backpropagation Improves Training for Sparse Mixture-of-ExpertsAshwinee Panda, Vatsal Baherwani, Zain Sarwar, Benjamin Thérien et al.NeurIPS 2025 · 10 citations
- ERMoE: Eigen-Reparameterized Mixture-of-Experts for Stable Routing and Interpretable SpecializationAnzhe Cheng, Shukai Duan, Shixuan Li, Chenzhong Yin et al.CVPR 2026 · 8 citations
- TA-MoE: Topology-Aware Large Scale Mixture-of-Expert TrainingChang Chen, Min Li, Zhihua Wu, Dianhai Yu et al.NeurIPS 2022 · 31 citations
- ReMoE: Fully Differentiable Mixture-of-Experts with ReLU RoutingZiteng Wang, Jun Zhu, Jianfei ChenICLR 2025
