Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing
Mengfan Liu, Wei Wang, Chuan Wu
摘要
With the advancement of serverless computing, running machine learning (ML) inference services over a serverless platform has been advocated, given its labor-free scalability and cost effectiveness. Mixture-of-Experts (MoE) models have been a dominant type of model architectures to enable large models nowadays, with parallel expert networks. Serving large MoE models on serverless computing is potentially beneficial, but has been underexplored due to substantial challenges in handling the skewed expert popularity and scatter-gather communication bottleneck in MoE model execution, for cost-efficient serverless MoE deployment and performance guarantee. We study optimized MoE model deployment and distributed inference serving on a serverless platform, that effectively predict expert selection, pipeline communication with model execution, and minimize the overall billed cost of serving MoE models. Especially, we propose a Bayesian optimization framework with multi-dimensionalsearch to learn expert selection and optimal MoE deployment achieving optimal billed cost, including: 1) a Bayesian decision-making method for predicting expert popularity; 2) flexibly pipelined scatter-gather communication; and 3) an optimal model deployment algorithm for distributed MoE serving. Extensive experiments on AWS Lambda show that our designs reduce the billed cost of all MoE layers by at least 75.67% compared to CPU clusters while maintaining satisfactory inference throughput. As compared to LambdaML in serverless computing, our design achieves 43.41% lower cost with a throughput decrease of no more than 18.76%.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Faster, Smaller, and Smarter: Task-Aware Expert Merging for Online MoE InferenceZiyi Han, Xutong Liu, Ruiting Zhou, Xiangxiang Dai 等INFOCOM 2026 · 被引用 3 次
- Sub-MoE: Efficient Mixture-of-Expert LLMs Compression via Subspace Expert MergingLujun Li, Qiyuan Zhu, Jiacheng Wang, Xiaoyu Qin 等AAAI 2026 · 被引用 2 次
- Taming Latency-Memory Trade-Off in MoE-Based LLM Serving via Fine-Grained Expert OffloadingHanfei Yu, Xingqi Cui, Hong Zhang, Hao Wang 等EuroSys 2026 · 被引用 1 次
它引用的顶会 Paper18
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen 等ICLR 2021 · 被引用 1,954 次
- GLaM: Efficient Scaling of Language Models with Mixture-of-ExpertsNan Du, Yanping Huang, Andrew M. Dai, Simon Tong 等ICML 2022 · 被引用 1,173 次
- Mixture-of-Experts with Expert Choice RoutingYanqi Zhou, Tao Lei, Hanxiao Liu, Nan Du 等NeurIPS 2022 · 被引用 933 次
- DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI ScaleSamyam Rajbhandari, Conglong Li, Zhewei Yao, Minjia Zhang 等ICML 2022 · 被引用 523 次
- INFaaS: Automated Model-less Inference ServingFrancisco Romero, Qian Li, Neeraja J. Yadwadkar, Christos KozyrakisUSENIX ATC 2021 · 被引用 325 次
相关 Paper
- ZipMoE: Efficient On-Device MoE Serving via Lossless Compression and Cache-Affinity SchedulingYuchen Yang, Yaru Zhao, Pu Yang, Shaowei Wang 等ICML 2026 · 被引用 3 次
- Dynamo-MoE: Accelerating Sparse Large Model Inference with Dynamic ParallelizationJiahao Chen, Shigang Li, Rongtian Fu, Tong Wu 等HPDC 2026
- Semantic Parallelism: Redefining Efficient MoE Inference via Model-Data Co-SchedulingYan Li, Zhenyu Zhang, Zhengang Wang, Pengfei chen 等ICLR 2026 · 被引用 11 次
- Accelerating Distributed MoE Training and Inference with LinaJiamin Li, Yimin Jiang, Yibo Zhu, Cong Wang 等USENIX ATC 2023 · 被引用 191 次
- MegaScale-Infer: Efficient Mixture-of-Experts Model Serving with Disaggregated Expert ParallelismRuidong Zhu, Ziheng Jiang, Chao Jin, Peng Wu 等SIGCOMM 2025 · 被引用 19 次
