Lune

ICLR2026Top-tier venue

Libra: Effective yet Efficient Load Balancing for Large-scale MoE Inference

Jaehoon Yang, Yushin Kim, Seokwon Moon, Yeonhong Park, Jae W. Lee

2026Year

Abstract

Distributed inference of large-scale Mixture-of-Experts (MoE) models faces a critical challenge: expert load imbalance. Numerous system-level approaches have been proposed for load balancing, but they either fail to achieve a satisfactory level of balance or introduce new bottlenecks due to the overhead of the load balancing mechanism itself. To this end, we propose Libra, a system that achieves near-optimal load balancing with minimal overhead. Libra adopts sophisticated mechanisms that accurately predict future expert activations and, based on these predictions, systematically perform load balancing. At the same time, it effectively hides the associated overhead by reconstructing the execution flow so that these costs are overlapped with MoE computation. Evaluations with two large-scale state-of-the-art MoE models on 8 H200 GPUs demonstrate that Libra improves throughput by up to 19.2%. The code is available at https://github.com/SNU-ARC/Libra.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 72609fd7-17cd-4462-a27c-94d05ccdee13

Builds on16

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines