CCLInsight: Unveiling Insights in GPU Collective Communication Libraries via Primitive-Centric Analysis
Liuyao Dai, Adam Weingram, Weicong Chen, Xiaoyi Lu
摘要
As distributed AI workloads scale in complexity, GPU-based collective communication libraries (xCCLs) such as NCCL are becoming increasingly critical for efficient data movement across heterogeneous hardware. Characterizing and optimizing the interactions between hardware architectures, communication algorithm designs, and low-level primitive configurations is essential to fully utilize GPU and interconnect resources while ensuring scalability. However, existing analysis methods often lack depth, limit optimization insights, leave significant performance gains untapped, or lack generality. To address these limitations, we propose a novel primitive-centric profiling and analysis method that integrates architectural, algorithmic, and primitive-level insights to comprehensively qualify and quantify performance-critical parameters in GPU-based CCLs. We implement and evaluate this method in our tool CCLInsight. Evaluating three large-scale GPU clusters with diverse GPU architectures (64 H100s, 64 RTX 5000s, and 256 A100s) for design analysis, and a fourth 256-GPU H200 cluster for application-level evaluation, CCLInsight uncovers critical parameter influences on communication efficiency and scalability, as well as exposing a new CCL performance scaling law, the collapsed optimal configuration scaling law. Adjusting parameters based on this analysis, we demonstrate a 43.90% improvement in NCCL and a 40.64× speedup in MSCCL over default configurations in the collective communication microbenchmark, while also improving NCCL and MSCCL performance by up to 2.27× and 3.00×, respectively, on LLM training workloads (GPT-3, LLaMA2, DeepSeek-R1). These contributions establish CCLInsight as a useful tool for xCCL characterization and optimization in distributed AI environments.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- TCCL: Discovering Better Communication Paths for PCIe GPU ClustersHeehoon Kim, Junyeol Ryu, Jaejin LeeASPLOS 2024 · 被引用 26 次
- SyCCL: Exploiting Symmetry for Efficient Collective Communication SchedulingJiamin Cao, Shangfeng Shi, Jiaqi Gao, Weisen Liu 等SIGCOMM 2025 · 被引用 15 次
- AutoCCL: Automated Collective Communication Tuning for Accelerating Distributed and Parallel DNN TrainingGuanbin Xu, Zhihao Le, Yinhe Chen, Zhiqi Lin 等NSDI 2025 · 被引用 27 次
- HeteroSim: Towards High-Fidelity Heterogeneous LLM Training Simulation on GPUsXiaofei Yue, Fangming Zhao, Fulun Ye, Jiongchi Yu 等WWW 2026
- PC4: Precision Collective Communication Congestion Control for AI ClusterTaoran Qi, Shuo Li, Xingqi Zou, Liangce Deng 等INFOCOM 2025 · 被引用 3 次
