FasterMoE: modeling and optimizing training of large-scale dynamic pre-trained models
Jiaao He, Jidong Zhai, Tiago Antunes, Haojie Wang, Fuwen Luo, Shangfeng Shi, Qin Li
摘要
The current trend in deep learning is to scale models to extremely large sizes with the objective of increasing their accuracy. Mixture-of-Expert (MoE) is the most popular pre-trained model that makes feasible the training of models with parameters beyond trillion-scale. Thanks to the dynamic activation of experts, i.e., shallow layers specialized in certain domains, it allows for sparse training of bigger models, removing the linearity between model size and computation. However, different from traditional deep learning models, it draws huge challenges to the efficiency of these training systems, including dynamic load imbalance, inefficient synchronous execution mode, and congested all-to-all communication.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper58
- Accelerating Distributed MoE Training and Inference with LinaJiamin Li, Yimin Jiang, Yibo Zhu, Cong Wang 等USENIX ATC 2023 · 被引用 191 次
- Merge, Then Compress: Demystify Efficient SMoE with Hints from Its Routing PolicyPingzhi Li, Zhenyu Zhang, Prateek Yadav, Yi-Lin Sung 等ICLR 2024 · 被引用 97 次
- SmartMoE: Efficiently Training Sparsely-Activated Models through Combining Offline and Online ParallelizationMingshu Zhai, Jiaao He, Zixuan Ma, Zan Zong 等USENIX ATC 2023 · 被引用 96 次
- NanoFlow: Towards Optimal Large Language Model Serving ThroughputKan Zhu, Yufei Gao, Yilong Zhao, Liangyu Zhao 等OSDI 2025 · 被引用 92 次
- Toward Efficient Inference for Mixture of ExpertsHaiyang Huang, Newsha Ardalani, Anna Y. Sun, Liu Ke 等NeurIPS 2024 · 被引用 60 次
相关 Paper
- Harnessing Inter-GPU Shared Memory for Seamless MoE Communication-Computation FusionHulin Wang, Yaqi Xia, Donglin Yang, Xiaobo Zhou 等PPoPP 2025 · 被引用 7 次
- Janus: A Unified Distributed Training Framework for Sparse Mixture-of-Experts ModelsJuncai Liu, Jessie Hui Wang, Yimin JiangSIGCOMM 2023 · 被引用 46 次
- MegaScale-MoE: Large-Scale Communication-Efficient Training of Mixture-of-Experts Models in ProductionChao Jin, Ziheng Jiang, Zhihao Bai, Zheng Zhong 等EuroSys 2026 · 被引用 5 次
- PopFetcher: Towards Accelerated Mixture-of-Experts Training Via Popularity Based Expert-Wise PrefetchJunyi Zhang, Chuanhu Ma, Xiong Wang, Yuntao Nie 等USENIX ATC 2025 · 被引用 8 次
- Scaling Beyond the GPU Memory Limit for Large Mixture-of-Experts Model TrainingYechan Kim, Hwijoon Lim, Dongsu HanICML 2024 · 被引用 10 次
