Lune

HPDC2026顶会

Dynamo-MoE: Accelerating Sparse Large Model Inference with Dynamic Parallelization

Jiahao Chen, Shigang Li, Rongtian Fu, Tong Wu, Zhi Ma, Jingkun Dong

2026年份

摘要

Mixtral-of-Experts (MoE) has become one of the major model structures in LLMs because of its computational efficiency when scaling the model size. However, MoE model inference suffers from critical load imbalance issue caused by the sparsely and dynamically activated experts. In addition, current inference frameworks are oblivious to the real-time workload fluctuation, a common phenomenon in LLM serving. Therefore, the static model deployment of existing frameworks leads to severe performance limitations. To this end, we propose Dynamo-MoE, an out-of-box MoE inference framework to bridge the performance gap by dynamic parallelization strategies. Specifically, Dynamo-MoE integrates a novel load balancing approach based on token sorting and on-demand expert loading to solve the workload imbalance issue in the scenario of high workload (such as Prefill). Dynamo-MoE is also aware of workload varying to adaptively switch between tensor parallelism (for low latency in small batch scenarios) and expert parallelism (for high throughput in large batch scenarios). Furthermore, the model parameter redistribution overhead of dynamic parallelization is smartly overlapped through sophisticated pipeline orchestration. Compared to the SOTA framework vLLM (w/ and w/o EPLB), Dynamo-MoE achieves up to 6.75 × reduction for TTFT, 1.59 × reduction for TPOT, and 1.5 × improvement for throughput.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

lune papers get 5f0e8e57-064e-4b63-be8d-326d37bd2a8a

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖