Occult: Optimizing Collaborative Communications across Experts for Accelerated Parallel MoE Training and Inference
Shuqing Luo, Pingzhi Li, Jie Peng, Yang Zhao, Yu Cao, Yu Cheng, Tianlong Chen
Abstract
Mixture-of-experts (MoE) architectures could achieve impressive computational efficiency with expert parallelism, which relies heavily on all-toall communication across devices. Unfortunately, such communication overhead typically constitutes a significant portion of the total runtime, hampering the scalability of distributed training and inference for modern MoE models (consuming over 40% runtime in large-scale training). In this paper, we first define collaborative communication to illustrate this intrinsic limitation, and then propose system-and algorithm-level innovations to reduce communication costs. Specifically, given a pair of experts co-activated by one token, we call them as collaborated, which comprises 2 cases as intraand inter-collaboration, depending on whether they are kept on the same device. Our pilot investigations reveal that augmenting the proportion of intra-collaboration can accelerate expert parallel at scale. It motivates us to strategically optimize collaborative communication for accelerated MoE training and inference, dubbed Occult. Our designs are capable of either delivering exact results with reduced communication cost, or controllably minimizing the cost with collaboration pruning, materialized by modified fine-tuning. Comprehensive experiments on various MoE-LLMs demonstrate that Occult can be faster than popular state-of-theart inference or training frameworks (more than 1.5× speed up across multiple tasks and models) with comparable or superior quality compared to the standard fine-tuning. Code is available at https://github.com/UNITES-Lab/Occult .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Mozart: Modularized and Efficient MoE Training on 3.5D Wafer-Scale Chiplet ArchitecturesShuqing Luo, Ye Han, Pingzhi Li, Jiayin Qin et al.NeurIPS 2025 · 3 citations
- SCHUR-A*: Layer-wise Optimal Expert Pruning for MoEs via Schur-Complement Guided A* SearchZheng Chen, Weifeng Yang, Jianxiao Tang, Buhui YaoICML 2026
- EasyBalance: Cross-Layer Load Balancing in Distributed MoE InferenceYize Wu, KE GAO, Ling Li, Yanjun WuICML 2026
Builds on11
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI ScaleSamyam Rajbhandari, Conglong Li, Zhewei Yao, Minjia Zhang et al.ICML 2022 · 523 citations
- Hash Layers For Large Sparse ModelsStephen Roller, Sainbayar Sukhbaatar, Arthur Szlam, Jason WestonNeurIPS 2021 · 316 citations
Related papers
- Shortcut-connected Expert Parallelism for Accelerating Mixture of ExpertsWeilin Cai, Juyong Jiang, Le Qin, Junwei Cui et al.ICML 2025
- MegaScale-MoE: Large-Scale Communication-Efficient Training of Mixture-of-Experts Models in ProductionChao Jin, Ziheng Jiang, Zhihao Bai, Zheng Zhong et al.EuroSys 2026 · 5 citations
- FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts ModelsXinglin Pan, Wenxiang Lin, Lin Zhang, Shaohuai Shi et al.ASPLOS 2025 · 12 citations
- FoldMoE: Efficient Long Sequence MoE Training via Attention-MoE PipeliningGuichao Zhu, Lintian Lei, Yuhao Qing, Yichao Fu et al.ACL 2025
- BigMac: A Communication-Efficient Mixture-of-Experts Model Structure for Fast Training and InferenceZewen Jin, Shengnan Wang, Jiaan Zhu, Hongrui Zhan et al.AAAI 2025 · 6 citations
