CCLInsight: Unveiling Insights in GPU Collective Communication Libraries via Primitive-Centric Analysis
Liuyao Dai, Adam Weingram, Weicong Chen, Xiaoyi Lu
Abstract
As distributed AI workloads scale in complexity, GPU-based collective communication libraries (xCCLs) such as NCCL are becoming increasingly critical for efficient data movement across heterogeneous hardware. Characterizing and optimizing the interactions between hardware architectures, communication algorithm designs, and low-level primitive configurations is essential to fully utilize GPU and interconnect resources while ensuring scalability. However, existing analysis methods often lack depth, limit optimization insights, leave significant performance gains untapped, or lack generality. To address these limitations, we propose a novel primitive-centric profiling and analysis method that integrates architectural, algorithmic, and primitive-level insights to comprehensively qualify and quantify performance-critical parameters in GPU-based CCLs. We implement and evaluate this method in our tool CCLInsight. Evaluating three large-scale GPU clusters with diverse GPU architectures (64 H100s, 64 RTX 5000s, and 256 A100s) for design analysis, and a fourth 256-GPU H200 cluster for application-level evaluation, CCLInsight uncovers critical parameter influences on communication efficiency and scalability, as well as exposing a new CCL performance scaling law, the collapsed optimal configuration scaling law. Adjusting parameters based on this analysis, we demonstrate a 43.90% improvement in NCCL and a 40.64× speedup in MSCCL over default configurations in the collective communication microbenchmark, while also improving NCCL and MSCCL performance by up to 2.27× and 3.00×, respectively, on LLM training workloads (GPT-3, LLaMA2, DeepSeek-R1). These contributions establish CCLInsight as a useful tool for xCCL characterization and optimization in distributed AI environments.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 53c3d0ca-3ef3-429f-8471-ab7002691e1fRelated papers
- TCCL: Discovering Better Communication Paths for PCIe GPU ClustersHeehoon Kim, Junyeol Ryu, Jaejin LeeASPLOS 2024 · 26 citations
- SyCCL: Exploiting Symmetry for Efficient Collective Communication SchedulingJiamin Cao, Shangfeng Shi, Jiaqi Gao, Weisen Liu et al.SIGCOMM 2025 · 15 citations
- AutoCCL: Automated Collective Communication Tuning for Accelerating Distributed and Parallel DNN TrainingGuanbin Xu, Zhihao Le, Yinhe Chen, Zhiqi Lin et al.NSDI 2025 · 27 citations
- HeteroSim: Towards High-Fidelity Heterogeneous LLM Training Simulation on GPUsXiaofei Yue, Fangming Zhao, Fulun Ye, Jiongchi Yu et al.WWW 2026
- PC4: Precision Collective Communication Congestion Control for AI ClusterTaoran Qi, Shuo Li, Xingqi Zou, Liangce Deng et al.INFOCOM 2025 · 3 citations
