Occult: Optimizing Collaborative Communications across Experts for Accelerated Parallel MoE Training and Inference
Shuqing Luo, Pingzhi Li, Jie Peng, Yang Zhao, Yu Cao, Yu Cheng, Tianlong Chen
摘要
Mixture-of-experts (MoE) architectures could achieve impressive computational efficiency with expert parallelism, which relies heavily on all-toall communication across devices. Unfortunately, such communication overhead typically constitutes a significant portion of the total runtime, hampering the scalability of distributed training and inference for modern MoE models (consuming over 40% runtime in large-scale training). In this paper, we first define collaborative communication to illustrate this intrinsic limitation, and then propose system-and algorithm-level innovations to reduce communication costs. Specifically, given a pair of experts co-activated by one token, we call them as collaborated, which comprises 2 cases as intraand inter-collaboration, depending on whether they are kept on the same device. Our pilot investigations reveal that augmenting the proportion of intra-collaboration can accelerate expert parallel at scale. It motivates us to strategically optimize collaborative communication for accelerated MoE training and inference, dubbed Occult. Our designs are capable of either delivering exact results with reduced communication cost, or controllably minimizing the cost with collaboration pruning, materialized by modified fine-tuning. Comprehensive experiments on various MoE-LLMs demonstrate that Occult can be faster than popular state-of-theart inference or training frameworks (more than 1.5× speed up across multiple tasks and models) with comparable or superior quality compared to the standard fine-tuning. Code is available at https://github.com/UNITES-Lab/Occult .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Mozart: Modularized and Efficient MoE Training on 3.5D Wafer-Scale Chiplet ArchitecturesShuqing Luo, Ye Han, Pingzhi Li, Jiayin Qin 等NeurIPS 2025 · 被引用 3 次
- SCHUR-A*: Layer-wise Optimal Expert Pruning for MoEs via Schur-Complement Guided A* SearchZheng Chen, Weifeng Yang, Jianxiao Tang, Buhui YaoICML 2026
- EasyBalance: Cross-Layer Load Balancing in Distributed MoE InferenceYize Wu, KE GAO, Ling Li, Yanjun WuICML 2026
它引用的顶会 Paper11
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI ScaleSamyam Rajbhandari, Conglong Li, Zhewei Yao, Minjia Zhang 等ICML 2022 · 被引用 523 次
- Hash Layers For Large Sparse ModelsStephen Roller, Sainbayar Sukhbaatar, Arthur Szlam, Jason WestonNeurIPS 2021 · 被引用 316 次
相关 Paper
- Shortcut-connected Expert Parallelism for Accelerating Mixture of ExpertsWeilin Cai, Juyong Jiang, Le Qin, Junwei Cui 等ICML 2025
- MegaScale-MoE: Large-Scale Communication-Efficient Training of Mixture-of-Experts Models in ProductionChao Jin, Ziheng Jiang, Zhihao Bai, Zheng Zhong 等EuroSys 2026 · 被引用 5 次
- FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts ModelsXinglin Pan, Wenxiang Lin, Lin Zhang, Shaohuai Shi 等ASPLOS 2025 · 被引用 12 次
- FoldMoE: Efficient Long Sequence MoE Training via Attention-MoE PipeliningGuichao Zhu, Lintian Lei, Yuhao Qing, Yichao Fu 等ACL 2025
- BigMac: A Communication-Efficient Mixture-of-Experts Model Structure for Fast Training and InferenceZewen Jin, Shengnan Wang, Jiaan Zhu, Hongrui Zhan 等AAAI 2025 · 被引用 6 次
