Libra: Effective yet Efficient Load Balancing for Large-scale MoE Inference
Jaehoon Yang, Yushin Kim, Seokwon Moon, Yeonhong Park, Jae W. Lee
Abstract
Distributed inference of large-scale Mixture-of-Experts (MoE) models faces a critical challenge: expert load imbalance. Numerous system-level approaches have been proposed for load balancing, but they either fail to achieve a satisfactory level of balance or introduce new bottlenecks due to the overhead of the load balancing mechanism itself. To this end, we propose Libra, a system that achieves near-optimal load balancing with minimal overhead. Libra adopts sophisticated mechanisms that accurately predict future expert activations and, based on these predictions, systematically perform load balancing. At the same time, it effectively hides the associated overhead by reconstructing the execution flow so that these costs are overlapped with MoE computation. Evaluations with two large-scale state-of-the-art MoE models on 8 H200 GPUs demonstrate that Libra improves throughput by up to 19.2%. The code is available at https://github.com/SNU-ARC/Libra.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 72609fd7-17cd-4462-a27c-94d05ccdee13Builds on16
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 2,317 citations
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen et al.ICLR 2021 · 1,954 citations
- GLaM: Efficient Scaling of Language Models with Mixture-of-ExpertsNan Du, Yanping Huang, Andrew M. Dai, Simon Tong et al.ICML 2022 · 1,173 citations
- DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model ServingYinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu et al.OSDI 2024 · 646 citations
- LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation DatasetLianmin Zheng, Wei-Lin Chiang, Ying Sheng, Tianle Li et al.ICLR 2024 · 419 citations
Related papers
- EasyBalance: Cross-Layer Load Balancing in Distributed MoE InferenceYize Wu, KE GAO, Ling Li, Yanjun WuICML 2026
- Toward Efficient Inference for Mixture of ExpertsHaiyang Huang, Newsha Ardalani, Anna Y. Sun, Liu Ke et al.NeurIPS 2024 · 60 citations
- FloE: On-the-Fly MoE Inference on Memory-constrained GPUYuxin Zhou, Zheng Li, Jun Zhang, Jue Wang et al.ICML 2025
- HybriMoE: Hybrid CPU-GPU Scheduling and Cache Management for Efficient MoE InferenceShuzhang Zhong, Yanfan Sun, Ling Liang, Runsheng Wang et al.DAC 2025 · 8 citations
- Harnessing Inter-GPU Shared Memory for Seamless MoE Communication-Computation FusionHulin Wang, Yaqi Xia, Donglin Yang, Xiaobo Zhou et al.PPoPP 2025 · 7 citations
