FasterMoE: modeling and optimizing training of large-scale dynamic pre-trained models
Jiaao He, Jidong Zhai, Tiago Antunes, Haojie Wang, Fuwen Luo, Shangfeng Shi, Qin Li
Abstract
The current trend in deep learning is to scale models to extremely large sizes with the objective of increasing their accuracy. Mixture-of-Expert (MoE) is the most popular pre-trained model that makes feasible the training of models with parameters beyond trillion-scale. Thanks to the dynamic activation of experts, i.e., shallow layers specialized in certain domains, it allows for sparse training of bigger models, removing the linearity between model size and computation. However, different from traditional deep learning models, it draws huge challenges to the efficiency of these training systems, including dynamic load imbalance, inefficient synchronous execution mode, and congested all-to-all communication.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 71a84ef9-a600-4020-9af4-be110d238512Cited by top-tier papers58
- Accelerating Distributed MoE Training and Inference with LinaJiamin Li, Yimin Jiang, Yibo Zhu, Cong Wang et al.USENIX ATC 2023 · 191 citations
- Merge, Then Compress: Demystify Efficient SMoE with Hints from Its Routing PolicyPingzhi Li, Zhenyu Zhang, Prateek Yadav, Yi-Lin Sung et al.ICLR 2024 · 97 citations
- SmartMoE: Efficiently Training Sparsely-Activated Models through Combining Offline and Online ParallelizationMingshu Zhai, Jiaao He, Zixuan Ma, Zan Zong et al.USENIX ATC 2023 · 96 citations
- NanoFlow: Towards Optimal Large Language Model Serving ThroughputKan Zhu, Yufei Gao, Yilong Zhao, Liangyu Zhao et al.OSDI 2025 · 92 citations
- Toward Efficient Inference for Mixture of ExpertsHaiyang Huang, Newsha Ardalani, Anna Y. Sun, Liu Ke et al.NeurIPS 2024 · 60 citations
Related papers
- Harnessing Inter-GPU Shared Memory for Seamless MoE Communication-Computation FusionHulin Wang, Yaqi Xia, Donglin Yang, Xiaobo Zhou et al.PPoPP 2025 · 7 citations
- Janus: A Unified Distributed Training Framework for Sparse Mixture-of-Experts ModelsJuncai Liu, Jessie Hui Wang, Yimin JiangSIGCOMM 2023 · 46 citations
- MegaScale-MoE: Large-Scale Communication-Efficient Training of Mixture-of-Experts Models in ProductionChao Jin, Ziheng Jiang, Zhihao Bai, Zheng Zhong et al.EuroSys 2026 · 5 citations
- PopFetcher: Towards Accelerated Mixture-of-Experts Training Via Popularity Based Expert-Wise PrefetchJunyi Zhang, Chuanhu Ma, Xiong Wang, Yuntao Nie et al.USENIX ATC 2025 · 8 citations
- Scaling Beyond the GPU Memory Limit for Large Mixture-of-Experts Model TrainingYechan Kim, Hwijoon Lim, Dongsu HanICML 2024 · 10 citations
