Libra: Effective yet Efficient Load Balancing for Large-scale MoE Inference
Jaehoon Yang, Yushin Kim, Seokwon Moon, Yeonhong Park, Jae W. Lee
摘要
Distributed inference of large-scale Mixture-of-Experts (MoE) models faces a critical challenge: expert load imbalance. Numerous system-level approaches have been proposed for load balancing, but they either fail to achieve a satisfactory level of balance or introduce new bottlenecks due to the overhead of the load balancing mechanism itself. To this end, we propose Libra, a system that achieves near-optimal load balancing with minimal overhead. Libra adopts sophisticated mechanisms that accurately predict future expert activations and, based on these predictions, systematically perform load balancing. At the same time, it effectively hides the associated overhead by reconstructing the execution flow so that these costs are overlapped with MoE computation. Evaluations with two large-scale state-of-the-art MoE models on 8 H200 GPUs demonstrate that Libra improves throughput by up to 19.2%. The code is available at https://github.com/SNU-ARC/Libra.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper16
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 被引用 2,317 次
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen 等ICLR 2021 · 被引用 1,954 次
- GLaM: Efficient Scaling of Language Models with Mixture-of-ExpertsNan Du, Yanping Huang, Andrew M. Dai, Simon Tong 等ICML 2022 · 被引用 1,173 次
- DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model ServingYinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu 等OSDI 2024 · 被引用 646 次
- LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation DatasetLianmin Zheng, Wei-Lin Chiang, Ying Sheng, Tianle Li 等ICLR 2024 · 被引用 419 次
相关 Paper
- EasyBalance: Cross-Layer Load Balancing in Distributed MoE InferenceYize Wu, KE GAO, Ling Li, Yanjun WuICML 2026
- Toward Efficient Inference for Mixture of ExpertsHaiyang Huang, Newsha Ardalani, Anna Y. Sun, Liu Ke 等NeurIPS 2024 · 被引用 60 次
- FloE: On-the-Fly MoE Inference on Memory-constrained GPUYuxin Zhou, Zheng Li, Jun Zhang, Jue Wang 等ICML 2025
- HybriMoE: Hybrid CPU-GPU Scheduling and Cache Management for Efficient MoE InferenceShuzhang Zhong, Yanfan Sun, Ling Liang, Runsheng Wang 等DAC 2025 · 被引用 8 次
- Harnessing Inter-GPU Shared Memory for Seamless MoE Communication-Computation FusionHulin Wang, Yaqi Xia, Donglin Yang, Xiaobo Zhou 等PPoPP 2025 · 被引用 7 次
