Scaling Beyond the GPU Memory Limit for Large Mixture-of-Experts Model Training
Yechan Kim, Hwijoon Lim, Dongsu Han
摘要
Mixture-of-Experts (MoE) is a powerful technique for enhancing the performance of neural networks while decoupling computational complexity from the number of parameters. However, despite this, scaling the number of experts requires adding more GPUs. In addition, the load imbalance in token load across experts causes unnecessary computation or straggler problems. We present ES-MoE, a novel method for efficient scaling MoE training. It offloads expert parameters to host memory and leverages pipelined expert processing to overlap GPU-CPU communication with GPU computation. It dynamically balances token loads across GPUs, improving computational efficiency. ES-MoE accelerates MoE training on a limited number of GPUs without degradation in model performance. We validate our approach on GPT-based MoE models, demonstrating 67× better scalability and up to 17.5× better throughput over existing frameworks. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Adaptive Multi-Scale Decomposition Framework for Time Series ForecastingYifan Hu, Peiyuan Liu, Peng Zhu, Dawei Cheng 等AAAI 2025 · 被引用 60 次
- HMoE: Heterogeneous Mixture of Experts for Language ModelingAn Wang, Xingwu Sun, Ruobing Xie, Shuaipeng Li 等EMNLP 2025 · 被引用 2 次
- Inner-layer Token Self-modulation as Another Scaling Axis for LLMsYebin Yang, Huaijin Wu, Jingtao Han, Yu Wang 等ICML 2026 · 被引用 1 次
- Achieving Cloud-Grade SLOs for Local Mixture-of-Experts Inference through CPU–GPU Hybrid DesignWenxin Wang, Yule Hou, Yu Ji, Peng Qu 等OSDI 2026 · 被引用 1 次
- Input Domain Aware MoE: Decoupling Routing Decisions from Task Optimization in Mixture of ExpertsYongXiang Hua, Haoyu Cao, Zhou Tao, Bocheng Li 等ACM MM 2025 · 被引用 1 次
它引用的顶会 Paper11
- Scaling Vision with Sparse Mixture of ExpertsCarlos Riquelme, Joan Puigcerver, Basil Mustafa, Maxim Neumann 等NeurIPS 2021 · 被引用 1,213 次
- GLaM: Efficient Scaling of Language Models with Mixture-of-ExpertsNan Du, Yanping Huang, Andrew M. Dai, Simon Tong 等ICML 2022 · 被引用 1,173 次
- Mixture-of-Experts with Expert Choice RoutingYanqi Zhou, Tao Lei, Hanxiao Liu, Nan Du 等NeurIPS 2022 · 被引用 933 次
- ZeRO-Offload: Democratizing Billion-Scale Model TrainingJie Ren, Samyam Rajbhandari, Reza Yazdani Aminabadi, Olatunji Ruwase 等USENIX ATC 2021 · 被引用 657 次
- DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI ScaleSamyam Rajbhandari, Conglong Li, Zhewei Yao, Minjia Zhang 等ICML 2022 · 被引用 523 次
相关 Paper
- X-MoE: Enabling Scalable Training for Emerging Mixture-of-Experts Architectures on HPC PlatformsYueming Yuan, Ahan Gupta, Jianping Li, Sajal Dash 等SC 2025 · 被引用 3 次
- LAER-MoE: Load-Adaptive Expert Re-layout for Efficient Mixture-of-Experts TrainingXinyi Liu, Yujie Wang, Fangcheng Fu, Xuefeng Xiao 等ASPLOS 2026
- Harnessing Inter-GPU Shared Memory for Seamless MoE Communication-Computation FusionHulin Wang, Yaqi Xia, Donglin Yang, Xiaobo Zhou 等PPoPP 2025 · 被引用 7 次
- SYMI: Efficient Mixture-of-Experts Training via Model and Optimizer State DecouplingAthinagoras Skiadopoulos, Mark Zhao, Swapnil Gandhi, Thomas Norrie 等NSDI 2026 · 被引用 6 次
- MegaScale-MoE: Large-Scale Communication-Efficient Training of Mixture-of-Experts Models in ProductionChao Jin, Ziheng Jiang, Zhihao Bai, Zheng Zhong 等EuroSys 2026 · 被引用 5 次
