Capacity-Aware Inference: Mitigating the Straggler Effect in Mixture of Experts
Shwai He, Weilin Cai, Jiayi Huang, Ang Li
Abstract
The Mixture of Experts (MoE) is an effective architecture for scaling large language models by leveraging sparse expert activation to balance performance and efficiency. However, under expert parallelism, MoE suffers from inference inefficiencies due to imbalanced token-to-expert assignment, where underloaded experts complete computations early but must wait for overloaded experts, leading to global delays. We define this phenomenon as the Straggler Effect, as the most burdened experts dictate the overall inference latency. To address this, we first propose Capacity-Aware Token Drop, which enforces expert capacity limits by discarding excess tokens from overloaded experts, effectively reducing load imbalance with minimal performance impact (e.g., speedup with only degradation on OLMoE).
Next, given the presence of low-load experts remaining well below the capacity threshold, we introduce Capacity-Aware Expanded Drop, which allows tokens to include additional local experts in their candidate set before enforcing strict local capacity constraints, thereby improving load balance and enhancing the utilization of underused experts.
Extensive experiments on both language and multimodal MoE models demonstrate the effectiveness of our approach, yielding substantial gains in expert utilization, model performance, and inference efficiency, e.g., applying Expanded Drop to Mixtral-87B-Instruct yields a 0.2% average performance improvement and a 1.85 inference speedup. The code is released at: https://github.com/CASE-Lab-UMD/Capacity-Aware-MoE.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Taming Latency-Memory Trade-Off in MoE-Based LLM Serving via Fine-Grained Expert OffloadingHanfei Yu, Xingqi Cui, Hong Zhang, Hao Wang et al.EuroSys 2026 · 1 citation
- RepetitionCurse: Measuring and Understanding Router Imbalance in Mixture-of-Experts LLMs under DoS StressRuixuan Huang, Qingyue Wang, Hantao Huang, Yudong Gao et al.ICML 2026
- Soft Modality-Guided Expert Specialization in MoE-VLMsZi-Hao Bo, Yaqian Li, Anzhou Hou, Rinyoichi Takezoe et al.CVPR 2026
- MACS: Modality-Aware Capacity Scaling for Efficient Multimodal MoE InferenceBo Li, Chuan Wu, Shaolin ZhuACL 2026
- Router-Tuning: A Simple and Effective Approach for Dynamic DepthShwai He, Tao Ge, Guoheng Sun, Bowei Tian et al.EMNLP 2025
Builds on13
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen et al.ICLR 2021 · 1,954 citations
- Scaling Vision with Sparse Mixture of ExpertsCarlos Riquelme, Joan Puigcerver, Basil Mustafa, Maxim Neumann et al.NeurIPS 2021 · 1,213 citations
- GLaM: Efficient Scaling of Language Models with Mixture-of-ExpertsNan Du, Yanping Huang, Andrew M. Dai, Simon Tong et al.ICML 2022 · 1,173 citations
- Mixture-of-Experts with Expert Choice RoutingYanqi Zhou, Tao Lei, Hanxiao Liu, Nan Du et al.NeurIPS 2022 · 933 citations
Related papers
- Mining Tensor/Neuron-Level Sparsity to Maximize Mixture-of-Experts Potential in Post-Training and InferenceWeilin Cai, Le Qin, Shwai He, Junwei Cui et al.ICML 2026
- Alloc-MoE: Budget-Aware Expert Activation Allocation for Efficient Mixture-of-Experts InferenceBaihui Liu, Kaiyuan Tian, Wei Wang, Zhaoning Zhang et al.ACL 2026
- Scaling Beyond the GPU Memory Limit for Large Mixture-of-Experts Model TrainingYechan Kim, Hwijoon Lim, Dongsu HanICML 2024 · 10 citations
- Dynamo-MoE: Accelerating Sparse Large Model Inference with Dynamic ParallelizationJiahao Chen, Shigang Li, Rongtian Fu, Tong Wu et al.HPDC 2026
- Klotski: Efficient Mixture-of-Expert Inference via Expert-Aware Multi-Batch PipelineZhiyuan Fang, Yuegui Huang, Zicong Hong, Yufeng Lyu et al.ASPLOS 2025 · 6 citations
