Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing
Mengfan Liu, Wei Wang, Chuan Wu
Abstract
With the advancement of serverless computing, running machine learning (ML) inference services over a serverless platform has been advocated, given its labor-free scalability and cost effectiveness. Mixture-of-Experts (MoE) models have been a dominant type of model architectures to enable large models nowadays, with parallel expert networks. Serving large MoE models on serverless computing is potentially beneficial, but has been underexplored due to substantial challenges in handling the skewed expert popularity and scatter-gather communication bottleneck in MoE model execution, for cost-efficient serverless MoE deployment and performance guarantee. We study optimized MoE model deployment and distributed inference serving on a serverless platform, that effectively predict expert selection, pipeline communication with model execution, and minimize the overall billed cost of serving MoE models. Especially, we propose a Bayesian optimization framework with multi-dimensionalsearch to learn expert selection and optimal MoE deployment achieving optimal billed cost, including: 1) a Bayesian decision-making method for predicting expert popularity; 2) flexibly pipelined scatter-gather communication; and 3) an optimal model deployment algorithm for distributed MoE serving. Extensive experiments on AWS Lambda show that our designs reduce the billed cost of all MoE layers by at least 75.67% compared to CPU clusters while maintaining satisfactory inference throughput. As compared to LambdaML in serverless computing, our design achieves 43.41% lower cost with a throughput decrease of no more than 18.76%.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Faster, Smaller, and Smarter: Task-Aware Expert Merging for Online MoE InferenceZiyi Han, Xutong Liu, Ruiting Zhou, Xiangxiang Dai et al.INFOCOM 2026 · 3 citations
- Sub-MoE: Efficient Mixture-of-Expert LLMs Compression via Subspace Expert MergingLujun Li, Qiyuan Zhu, Jiacheng Wang, Xiaoyu Qin et al.AAAI 2026 · 2 citations
- Taming Latency-Memory Trade-Off in MoE-Based LLM Serving via Fine-Grained Expert OffloadingHanfei Yu, Xingqi Cui, Hong Zhang, Hao Wang et al.EuroSys 2026 · 1 citation
Builds on18
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen et al.ICLR 2021 · 1,954 citations
- GLaM: Efficient Scaling of Language Models with Mixture-of-ExpertsNan Du, Yanping Huang, Andrew M. Dai, Simon Tong et al.ICML 2022 · 1,173 citations
- Mixture-of-Experts with Expert Choice RoutingYanqi Zhou, Tao Lei, Hanxiao Liu, Nan Du et al.NeurIPS 2022 · 933 citations
- DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI ScaleSamyam Rajbhandari, Conglong Li, Zhewei Yao, Minjia Zhang et al.ICML 2022 · 523 citations
- INFaaS: Automated Model-less Inference ServingFrancisco Romero, Qian Li, Neeraja J. Yadwadkar, Christos KozyrakisUSENIX ATC 2021 · 325 citations
Related papers
- ZipMoE: Efficient On-Device MoE Serving via Lossless Compression and Cache-Affinity SchedulingYuchen Yang, Yaru Zhao, Pu Yang, Shaowei Wang et al.ICML 2026 · 3 citations
- Dynamo-MoE: Accelerating Sparse Large Model Inference with Dynamic ParallelizationJiahao Chen, Shigang Li, Rongtian Fu, Tong Wu et al.HPDC 2026
- Semantic Parallelism: Redefining Efficient MoE Inference via Model-Data Co-SchedulingYan Li, Zhenyu Zhang, Zhengang Wang, Pengfei chen et al.ICLR 2026 · 11 citations
- Accelerating Distributed MoE Training and Inference with LinaJiamin Li, Yimin Jiang, Yibo Zhu, Cong Wang et al.USENIX ATC 2023 · 191 citations
- MegaScale-Infer: Efficient Mixture-of-Experts Model Serving with Disaggregated Expert ParallelismRuidong Zhu, Ziheng Jiang, Chao Jin, Peng Wu et al.SIGCOMM 2025 · 19 citations
