AutoCCL: Automated Collective Communication Tuning for Accelerating Distributed and Parallel DNN Training
Guanbin Xu, Zhihao Le, Yinhe Chen, Zhiqi Lin, Zewen Jin, Youshan Miao, Cheng Li
Abstract
The collective communication libraries are pivotal in optimizing the performance of distributed and parallel deep neural network (DNN) training. Most network optimizations are under the assumption that these libraries are well-tuned, ignoring their low-level parameter selection. In this paper, we present a novel automated tuning method AutoCCL that significantly improves communication performance without incurring additional costs. One of the primary challenges we tackle is the state explosion in searching for the optimal configuration. To overcome this, we decouple implementation-related parameters from those sensitive to the search space size and propose a divide-and-conquer algorithm, minimizing the requirement for exhaustive trials. We further propose an online tuning approach that accounts for communication-computation interference to enhance accuracy in finding optimal configurations, while hiding tuning overhead within early iterations of training jobs. We implement AutoCCL atop NCCL, a leading and widely-used communication library provided by NVIDIA. Our evaluation on both a 2-node cluster (16 A40 GPUs, intranode NVLink, inter-node 2× 400Gbps InfiniBand) and a 4-node cluster (32 A40 GPUs, intra-node PCIe, inter-node 100Gbps InfiniBand) demonstrates that AutoCCL achieves 1.24-1.29× and 1.15-1.22× speedups on microbenchmarks compared to NCCL and another SOTA NCCL tuner, respectively, and up to 1.80× and 1.49× with concurrent computation. End-to-end evaluations on three large language models and one vision model show 1.07-1.32× improvements in periteration training time.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 11dcb6de-9c6d-4a16-862a-9502da3ecc2eCited by top-tier papers4
- SyCCL: Exploiting Symmetry for Efficient Collective Communication SchedulingJiamin Cao, Shangfeng Shi, Jiaqi Gao, Weisen Liu et al.SIGCOMM 2025 · 15 citations
- SwiftEP: Accelerating MoE Inference with Buffer Fusion and TMA OffloadingXingyi Li, Yadong Liu, Xiaojie Huang, Yiran Zhang et al.NSDI 2026 · 2 citations
- UCCL-Tran: An Extensible Software Transport Layer for GPU NetworkingYang Zhou, Zhongjie Chen, Ziming Mao, ChonLam Lao et al.OSDI 2026
- Lit Silicon: A Case Where Thermal Imbalance Couples Concurrent Execution in Multiple GPUsMarco Kurzynski, Shaizeen Aga, Di WuISCA 2026
Builds on29
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen et al.ICLR 2021 · 1,954 citations
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 852 citations
- Efficient large-scale language model training on GPU clusters using megatron-LMDeepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley et al.SC 2021 · 576 citations
- DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI ScaleSamyam Rajbhandari, Conglong Li, Zhewei Yao, Minjia Zhang et al.ICML 2022 · 523 citations
Related papers
- COCCL: A Collective Communication Library Supporting Easy Integration and Configuration of Customized Compression for Scalable LLM TrainingXingchen Liu, Haoran Kong, Hairui Zhao, Shengkai Lyu et al.PPoPP 2026 · 3 citations
- TCCL: Discovering Better Communication Paths for PCIe GPU ClustersHeehoon Kim, Junyeol Ryu, Jaejin LeeASPLOS 2024 · 26 citations
- TACCL: Guiding Collective Algorithm Synthesis using Communication SketchesAashaka Shah, Vijay Chidambaram, Meghan Cowan, Saeed Maleki et al.NSDI 2023
- xCCLTuner: Treating xCCL as Black-Box and Automatically TuningChenxu Wang, Wentao Fan, Zhehao Lin, Peirui Cao et al.INFOCOM 2026
- CCLInsight: Unveiling Insights in GPU Collective Communication Libraries via Primitive-Centric AnalysisLiuyao Dai, Adam Weingram, Weicong Chen, Xiaoyi LuICSE 2026 · 1 citation
