Toward Efficient Inference for Mixture of Experts
Haiyang Huang, Newsha Ardalani, Anna Y. Sun, Liu Ke, Shruti Bhosale, Hsien-Hsin S. Lee, Carole-Jean Wu, Benjamin Lee
摘要
Mixture-of-Experts (MoE) models have recently gained steam in achieving the state-of-the-art performance in a wide range of tasks in computer vision and natural language processing. They effectively expand the model capacity while incurring a minimal increase in computation cost during training. However, deploying such models for inference is difficult due to their large model size and complex communication pattern. In this work, we provide a characterization of two MoE workloads, namely Language Modeling (LM) and Machine Translation (MT) and identify their sources of inefficiencies at deployment. We propose three optimization techniques to mitigate sources of inefficiencies, namely (1) Dynamic gating, (2) Expert Buffering, and (3) Expert load balancing. We show that dynamic gating improves maximum throughput by 6.21-11.55 × for LM, 5.75-10.98 × for MT Encoder and 2.58-5.71 × for MT Decoder. It also reduces memory usage by up to 1.36 × for LM and up to 1.1 × for MT. We further propose Expert Buffering, a new caching mechanism that only keeps hot, active experts in GPU memory while buffering the rest in CPU memory. This reduces static memory allocation by 1.47 × . Finally, we propose a load balancing methodology that provides additional robustness to the workload. Our code is available at https://github.com/ hyhuang00/moe_inference .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- POD-Attention: Unlocking Full Prefill-Decode Overlap for Faster LLM InferenceAditya K. Kamath, Ramya Prabhu, Jayashree Mohan, Simon Peter 等ASPLOS 2025 · 被引用 24 次
- STEM: Scaling Transformers with Embedding ModulesRanajoy Sadhukhan, Sheng Cao, Harry Dong, Changsheng Zhao 等ICLR 2026 · 被引用 14 次
- Sparse Models, Sparse Safety: Unsafe Routes in Mixture-of-Experts LLMsYukun Jiang, Hai Huang, Mingjie Li, Yage Zhang 等ICML 2026 · 被引用 9 次
- In-Storage Acceleration of Retrieval Augmented Generation as a ServiceRohan Mahapatra, Harsha Santhanam, Christopher Priebe, Hanyang Xu 等ISCA 2025 · 被引用 9 次
- Least-Loaded Expert Parallelism: Load Balancing An Imbalanced Mixture-of-ExpertsXuan-Phi Nguyen, Shrey Pandit, Austin Xu, Caiming Xiong 等ICML 2026 · 被引用 5 次
它引用的顶会 Paper17
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 被引用 2,878 次
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen 等ICLR 2021 · 被引用 1,954 次
- Scaling Vision with Sparse Mixture of ExpertsCarlos Riquelme, Joan Puigcerver, Basil Mustafa, Maxim Neumann 等NeurIPS 2021 · 被引用 1,213 次
- GLaM: Efficient Scaling of Language Models with Mixture-of-ExpertsNan Du, Yanping Huang, Andrew M. Dai, Simon Tong 等ICML 2022 · 被引用 1,173 次
相关 Paper
- SMoE: An Algorithm-System Co-Design for Pushing MoE to the Edge via Expert SubstitutionGuoying Zhu, Meng Li, Haipeng Dai, Xuechen Liu 等ISCA 2026 · 被引用 4 次
- MoNDE: Mixture of Near-Data Experts for Large-Scale Sparse ModelsTaehyun Kim, Kwanseok Choi, Youngmock Cho, Jaehoon Cho 等DAC 2024 · 被引用 11 次
- Diff-MoE: Efficient Batched MoE Inference with Priority-Driven Differential Expert CachingKexin Li, Wenkan Huang, Qinggang Wang, Long Zheng 等SC 2025 · 被引用 3 次
- Pre-gated MoE: An Algorithm-System Co-Design for Fast and Scalable Mixture-of-Expert InferenceRanggi Hwang, Jianyu Wei, Shijie Cao, Changho Hwang 等ISCA 2024 · 被引用 48 次
- MoE-APEX: An Efficient MoE Inference System with Adaptive Precision Expert OffloadingPeng Tang, Jiacheng Liu, Xiaofeng Hou, Yifei Pu 等ASPLOS 2026 · 被引用 4 次
