MiCS: Near-linear Scaling for Training Gigantic Model on Public Cloud
Zhen Zhang, Shuai Zheng, Yida Wang, Justin Chiu, George Karypis, Trishul Chilimbi, Mu Li, Xin Jin
Abstract
Existing general purpose frameworks for gigantic model training, i.e., dense models with billions of parameters, cannot scale efficiently on cloud environment with various networking conditions due to large communication overheads. In this paper, we propose MiCS, which Minimizes the Communication Scale to bring down communication overhead. Specifically, by decreasing the number of participants in a communication collective, MiCS can utilize heterogeneous network bandwidth, reduce network traffic over slower links, reduce the latency of communications for maintaining high network bandwidth utilization, and amortize expensive global gradient synchronization overhead. Our evaluation on AWS shows that the system throughput of MiCS is up to 2.89× that of the state-of-the-art large model training systems. MiCS achieves near-linear scaling efficiency, which is up to 1.27× that of DeepSpeed. MiCS allows us to train a proprietary model with 100 billion parameters on 512 GPUs with 99.4% weak-scaling efficiency, and it is able to saturate over 54.5% theoretical computation power of each GPU on a public cloud with less GPU memory and more restricted networks than DGX-A100 clusters.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f5b0149f-6b1b-4aad-a157-fdfda82d357fCited by top-tier papers8
- How Large Language Models Will Disrupt Data ManagementRaul Castro Fernandez, Aaron J. Elmore, Michael J. Franklin, Sanjay Krishnan et al.VLDB 2023 · 127 citations
- Alpa: Automating Inter- and Intra-Operator Parallelism for Distributed Deep LearningLianmin Zheng, Zhuohan Li, Hao Zhang, Yonghao Zhuang et al.OSDI 2022 · 75 citations
- Flash-LLM: Enabling Low-Cost and Highly-Efficient Large Generative Model Inference With Unstructured SparsityHaojun Xia, Zhen Zheng, Yuchao Li, Donglin Zhuang et al.VLDB 2024 · 29 citations
- Slapo: A Schedule Language for Progressive Optimization of Large Deep Learning Model TrainingHongzheng Chen, Cody Hao Yu, Shuai Zheng, Zhen Zhang et al.ASPLOS 2024 · 8 citations
- Rethinking Memory and Communication Costs for Efficient Data Parallel Training of Large Language ModelsHanxiao Zhang, Lin Ju, Chan Wu, Jinjing Huang et al.NeurIPS 2024 · 4 citations
Builds on13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen et al.ICLR 2021 · 1,954 citations
- CoAtNet: Marrying Convolution and Attention for All Data SizesZihang Dai, Hanxiao Liu, Quoc V. Le, Mingxing TanNeurIPS 2021 · 1,747 citations
- Scaling Vision TransformersXiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, Lucas BeyerCVPR 2022 · 767 citations
- ZeRO-Offload: Democratizing Billion-Scale Model TrainingJie Ren, Samyam Rajbhandari, Reza Yazdani Aminabadi, Olatunji Ruwase et al.USENIX ATC 2021 · 657 citations
Related papers
- MAD-Max Beyond Single-Node: Enabling Large Machine Learning Model Acceleration on Distributed SystemsSamuel Hsia, Alicia Golden, Bilge Acun, Newsha Ardalani et al.ISCA 2024 · 8 citations
- JABAS: Joint Adaptive Batching and Automatic Scaling for DNN Training on Heterogeneous GPUsGyeongchan Yun, Junesoo Kang, Hyunjoon Jeong, Sanghyeon Eom et al.EuroSys 2025 · 2 citations
- PipeComm: Maximizing Link Utilization Through Pipeline-Aware Collective Communication SynthesisRuifan Xu, Yuze Luo, Yuhao Meng, Size Zheng et al.ISCA 2026
- Communication-Efficient Distributed Deep Learning with Merged Gradient Sparsification on GPUsShaohuai Shi, Qiang Wang, Xiaowen Chu, Bo Li et al.INFOCOM 2020 · 66 citations
- Herring: rethinking the parameter server at scale for the cloudIndu Thangakrishnan, Derya Cavdar, Can Karakus, Piyush Ghai et al.SC 2020 · 11 citations
