CO2: Efficient Distributed Training with Full Communication-Computation Overlap
Weigao Sun, Zhen Qin, Weixuan Sun, Shidi Li, Dong Li, Xuyang Shen, Yu Qiao, Yiran Zhong
Abstract
The fundamental success of large language models hinges upon the efficacious implementation of large-scale distributed training techniques. Nevertheless, building a vast, high-performance cluster featuring high-speed communication interconnectivity is prohibitively costly, and accessible only to prominent entities. In this work, we aim to lower this barrier and democratize large-scale training with limited bandwidth clusters. We propose a new approach called CO2 that introduces local-updating and asynchronous communication to the distributed data-parallel training, thereby facilitating the full overlap of COmunication with COmputation. CO2 is able to attain a high scalability even on extensive multi-node clusters constrained by very limited communication bandwidth. We further propose the staleness gap penalty and outer momentum clipping techniques together with CO2 to bolster its convergence and training stability. Besides, CO2 exhibits seamless integration with well-established ZeRO-series optimizers which mitigate memory consumption of model states with large model training. We also provide a mathematical proof of convergence, accompanied by the establishment of a stringent upper bound. Furthermore, we validate our findings through an extensive set of practical experiments encompassing a wide range of tasks in the fields of computer vision and natural language processing. These experiments serve to demonstrate the capabilities of CO2 in terms of convergence, generalization, and scalability when deployed across configurations comprising up to 128 A100 GPUs. The outcomes emphasize the outstanding capacity of CO2 to hugely improve scalability, no matter on clusters with 800Gbps RDMA or 80Gbps TCP/IP inter-node connections.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2181fb2d-04b7-475c-a5a5-b5884aa8f450Cited by top-tier papers3
- FlashMoE: Fast Distributed MoE in a Single KernelOsayamen Jonathan Aimuyo, Byungsoo Oh, Rachee SinghNeurIPS 2025 · 22 citations
- ACCO: Accumulate While You Communicate for Communication-Overlapped Sharded LLM TrainingAdel Nabli, Louis Fournier, Pierre Erbacher, Louis Serrano et al.NeurIPS 2025 · 5 citations
- Sequence Accumulation and Beyond: Infinite Context Length on Single GPU and Large ClustersWeigao Sun, Yongtuo Liu, Xiaqiang Tang, Xiaoyu MoAAAI 2025 · 3 citations
Builds on16
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 852 citations
- Don't Use Large Mini-batches, Use Local SGDTao Lin, Sebastian U. Stich, Kumar Kshitij Patel, Martin JaggiICLR 2020 · 462 citations
- cosFormer: Rethinking Softmax In AttentionZhen Qin, Weixuan Sun, Hui Deng, Dongxu Li et al.ICLR 2022 · 303 citations
- SlowMo: Improving Communication-Efficient Distributed SGD with Slow MomentumJianyu Wang, Vinayak Tantia, Nicolas Ballas, Michael G. RabbatICLR 2020 · 220 citations
Related papers
- DES-LOC: Desynced Low Communication Adaptive Optimizers for Foundation ModelsAlex Iacob, Lorenzo Sani, Mher Safaryan, Paris Giampouras et al.ICLR 2026 · 2 citations
- Factored Gossip DiLoCo: Reducing Blocking Communication within DiLoCoChamin Hewa Koneputugodage, Thalaiyasingam Ajanthan, Sameera Ramasinghe, Hadi Mohaghegh Dolatabadi et al.ICML 2026
- HALoS: Hierarchical Asynchronous Local SGD over Slow Networks for Geo-Distributed Large Language Model TrainingGeon-Woo Kim, Junbo Li, Shashidhar Gandham, Omar Baldonado et al.ICML 2025
- MT-DAO: Multi-Timescale Distributed Adaptive Optimizers with Local UpdatesAlex Iacob, Andrej Jovanovic, Mher Safaryan, Meghdad Kurmanji et al.ICLR 2026 · 4 citations
- COCCL: A Collective Communication Library Supporting Easy Integration and Configuration of Customized Compression for Scalable LLM TrainingXingchen Liu, Haoran Kong, Hairui Zhao, Shengkai Lyu et al.PPoPP 2026 · 3 citations
