Lune

ICML2026Top-tier venue

Multi-Head LatentMoE and Head Parallel: Communication-Efficient and Deterministic MoE Parallelism

Chenwei Cui, Rockwell Jackson, Benjamin Joseph Herrera, Ana Tarano, Hannah Kerner

2026Year
1Citations

Abstract

Large language models have transformed many applications but remain expensive to train. Sparse Mixture of Experts (MoE) addresses this through conditional computation, with Expert Parallel (EP) as the standard distributed training method. However, EP has three limitations: communication cost grows linearly with the number of activated experts k, load imbalance affects latency and memory usage, and data-dependent communication requires metadata exchange. We propose Multi-Head LatentMoE and Head Parallel (HP), a new architecture and parallelism achieving O(1) communication cost regardless of k, completely balanced traffic, and deterministic communication, all while remaining compatible with EP. To accelerate Multi-Head LatentMoE, we propose IO-aware routing and expert computation. Compared to MoE with EP, Multi-Head LatentMoE with HP trains up to 1.61× faster while having identical performance. With doubled granularity, it achieves higher overall performance while still being 1.11× faster. Our method makes multi-billion-parameter foundation model research more accessible. https://github.com/kerner-lab/ Sparse-GPT-Pretraining

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext d1896ac7-4c9c-4126-83b7-c27f9efa8ce0

Builds on6

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines