SC2021Top-tier venue
DistGNN: scalable distributed training for large-scale graph neural networks
Md. Vasimuddin, Sanchit Misra, Guixiang Ma, Ramanarayan Mohanty, Evangelos Georganas, Alexander Heinecke, Dhiraj D. Kalamkar, Nesreen K. Ahmed, Sasikanth Avancha
Abstract
Full-batch training on Graph Neural Networks (GNN) to learn the structure of large graphs is a critical problem that needs to scale to hundreds of compute nodes to be feasible. It is challenging due to large memory capacity and bandwidth requirements on a single compute node and high communication volumes across multiple nodes. In this paper, we present DistGNN that optimizes the well-known Deep Graph Library (DGL) for full-batch training on CPU clusters via an efficient shared memory implementation, communication reduction using a minimum vertex-cut graph partitioning algorithm and communication avoidance using a family of delayed-update algorithms. Our results on four common GNN benchmark datasets: Reddit, OGB-Products, OGB-Papers and Proteins, show up to 3.7× speed-up using a single CPU socket and up to 97× speed-up using 128 CPU sockets, respectively, over baseline DGL implementations running on a single CPU socket.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cca3602d-c4d9-4d52-9bd8-c36fb8405a72Cited by top-tier papers25
- NeutronStar: Distributed GNN Training with Hybrid Dependency ManagementQiange Wang, Yanfeng Zhang, Hao Wang, Chaoyi Chen et al.SIGMOD 2022 · 60 citations
- Betty: Enabling Large-Scale GNN Training with Batch-Level Graph PartitioningShuangyan Yang, Minjia Zhang, Wenqian Dong, Dong LiASPLOS 2023 · 43 citations
- Learn Locally, Correct Globally: A Distributed Algorithm for Training Graph Neural NetworksMorteza Ramezani, Weilin Cong, Mehrdad Mahdavi, Mahmut T. Kandemir et al.ICLR 2022 · 35 citations
- Graphite: optimizing graph neural networks on CPUs through cooperative software-hardware techniquesZhangxiaowen Gong, Houxiang Ji, Yao Yao, Christopher W. Fletcher et al.ISCA 2022 · 30 citations
- HongTu: Scalable Full-Graph GNN Training on Multiple GPUsQiange Wang, Yao Chen, Weng-Fai Wong, Bingsheng HeSIGMOD 2024 · 24 citations
Builds on3
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong et al.NeurIPS 2020 · 3,935 citations
- Reducing communication in graph neural network trainingAlok Tripathy, Katherine A. Yelick, Aydin BuluçSC 2020 · 67 citations
- FeatGraph: a flexible and efficient backend for graph neural network systemsYuwei Hu, Zihao Ye, Minjie Wang, Jiali Yu et al.SC 2020 · 57 citations
Related papers
- DGCL: an efficient communication library for distributed GNN trainingZhenkun Cai, Xiao Yan, Yidi Wu, Kaihao Ma et al.EuroSys 2021 · 103 citations
- WholeGraph: A Fast Graph Neural Network Training Framework with Multi-GPU Distributed Shared Memory ArchitectureDongxu Yang, Junhong Liu, Jiaxing Qi, Junjie LaiSC 2022 · 12 citations
- Efficient scaling of dynamic graph neural networksVenkatesan T. Chakaravarthy, Shivmaran S. Pandian, Saurabh Raje, Yogish Sabharwal et al.SC 2021 · 35 citations
- ParGNN: A Scalable Graph Neural Network Training Framework on multi-GPUsJunyu Gu, Shunde Li, Rongqiang Cao, Jue Wang et al.DAC 2025
- Optimizing Task Placement and Online Scheduling for Distributed GNN Training AccelerationZiyue Luo, Yixin Bao, Chuan WuINFOCOM 2022 · 12 citations
