Demons in the Detail: On Implementing Load Balancing Loss for Training Specialized Mixture-of-Expert Models
Zihan Qiu, Zeyu Huang, Bo Zheng, Kaiyue Wen, Zekun Wang, Rui Men, Ivan Titov, Dayiheng Liu, Jingren Zhou, Junyang Lin
摘要
This paper revisits the implementation of oad-alancing oss (LBL) when training Mixture-of-Experts (MoEs) models. Specifically, LBL for MoEs is defined as , where is the total number of experts, represents the frequency of expert being selected, and denotes the average gating score of the expert . Existing MoE training frameworks usually employ the parallel training strategy so that and the LBL are calculated within a and then averaged across parallel groups. In essence, a micro-batch for training billion-scale LLMs normally contains very few sequences. So, the micro-batch LBL is almost at the sequence level, and the router is pushed to distribute the token evenly within each sequence. Under this strict constraint, even tokens from a domain-specific sequence (, code) are uniformly routed to all experts, thereby inhibiting expert specialization. In this work, we propose calculating LBL using a to loose this constraint. Because a global-batch contains much more diverse sequences than a micro-batch, which will encourage load balance at the corpus level. Specifically, we introduce an extra communication step to synchronize across micro-batches and then use it to calculate the LBL. Through experiments on training MoEs-based LLMs (up to total parameters and tokens), we surprisingly find that the global-batch LBL strategy yields excellent performance gains in both pre-training perplexity and downstream tasks. Our analysis reveals that the global-batch LBL also greatly improves the domain specialization of MoE experts.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-FreeZihan Qiu, Zekun Wang, Bo Zheng, Zeyu Huang 等NeurIPS 2025 · 被引用 336 次
- ShinkaEvolve: Towards Open-Ended and Sample-Efficient Program EvolutionRobert T. Lange, Yuki Imajuku, Edoardo CetinICLR 2026 · 被引用 162 次
- Controlled LLM Training on Spectral SphereTian Xie, Haoming Luo, Haoyu Tang, Hu Yiwen 等ICML 2026 · 被引用 22 次
- MoRE: 3D Visual Geometry Reconstruction Meets Mixture-of-ExpertsJingnan Gao, Zhe Wang, Xianze Fang, Xingyu Ren 等CVPR 2026 · 被引用 19 次
- ProPhy: Progressive Physical Alignment for Dynamic World SimulationZijun Wang, Panwen Hu, Jing Wang, Terry Jingchen Zhang 等CVPR 2026 · 被引用 14 次
它引用的顶会 Paper8
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Mixture-of-Experts with Expert Choice RoutingYanqi Zhou, Tao Lei, Hanxiao Liu, Nan Du 等NeurIPS 2022 · 被引用 933 次
- DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI ScaleSamyam Rajbhandari, Conglong Li, Zhewei Yao, Minjia Zhang 等ICML 2022 · 被引用 523 次
- OpenMoE: An Early Effort on Open Mixture-of-Experts Language ModelsFuzhao Xue, Zian Zheng, Yao Fu, Jinjie Ni 等ICML 2024 · 被引用 183 次
- DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language ModelsDamai Dai, Chengqi Deng, Chenggang Zhao, R. X. Xu 等ACL 2024 · 被引用 171 次
相关 Paper
- Advancing Expert Specialization for Better MoEHongcan Guo, Haolang Lu, Guoshun Nan, Bolun Chu 等NeurIPS 2025 · 被引用 48 次
- Hierarchical Mixture of Experts with Two-Stage OptimizationGleb Molodtsov, Alexander Miasnikov, Aleksandr BeznosikovKDD 2026 · 被引用 2 次
- Coupling Experts and Routers in Mixture-of-Experts via an Auxiliary LossAng Lv, Jin Ma, Yiyuan Ma, Siyuan QiaoICLR 2026 · 被引用 14 次
- FoldMoE: Efficient Long Sequence MoE Training via Attention-MoE PipeliningGuichao Zhu, Lintian Lei, Yuhao Qing, Yichao Fu 等ACL 2025
- Not All Models Suit Expert Offloading: On Local Routing Consistency of Mixture-of-Expert ModelsJingcong Liang, Siyuan Wang, Miren Tian, Yitong Li 等ICLR 2026 · 被引用 8 次
