DistGNN: scalable distributed training for large-scale graph neural networks
Md. Vasimuddin, Sanchit Misra, Guixiang Ma, Ramanarayan Mohanty, Evangelos Georganas, Alexander Heinecke, Dhiraj D. Kalamkar, Nesreen K. Ahmed, Sasikanth Avancha
摘要
Full-batch training on Graph Neural Networks (GNN) to learn the structure of large graphs is a critical problem that needs to scale to hundreds of compute nodes to be feasible. It is challenging due to large memory capacity and bandwidth requirements on a single compute node and high communication volumes across multiple nodes. In this paper, we present DistGNN that optimizes the well-known Deep Graph Library (DGL) for full-batch training on CPU clusters via an efficient shared memory implementation, communication reduction using a minimum vertex-cut graph partitioning algorithm and communication avoidance using a family of delayed-update algorithms. Our results on four common GNN benchmark datasets: Reddit, OGB-Products, OGB-Papers and Proteins, show up to 3.7× speed-up using a single CPU socket and up to 97× speed-up using 128 CPU sockets, respectively, over baseline DGL implementations running on a single CPU socket.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper25
- NeutronStar: Distributed GNN Training with Hybrid Dependency ManagementQiange Wang, Yanfeng Zhang, Hao Wang, Chaoyi Chen 等SIGMOD 2022 · 被引用 60 次
- Betty: Enabling Large-Scale GNN Training with Batch-Level Graph PartitioningShuangyan Yang, Minjia Zhang, Wenqian Dong, Dong LiASPLOS 2023 · 被引用 43 次
- Learn Locally, Correct Globally: A Distributed Algorithm for Training Graph Neural NetworksMorteza Ramezani, Weilin Cong, Mehrdad Mahdavi, Mahmut T. Kandemir 等ICLR 2022 · 被引用 35 次
- Graphite: optimizing graph neural networks on CPUs through cooperative software-hardware techniquesZhangxiaowen Gong, Houxiang Ji, Yao Yao, Christopher W. Fletcher 等ISCA 2022 · 被引用 30 次
- HongTu: Scalable Full-Graph GNN Training on Multiple GPUsQiange Wang, Yao Chen, Weng-Fai Wong, Bingsheng HeSIGMOD 2024 · 被引用 24 次
它引用的顶会 Paper3
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong 等NeurIPS 2020 · 被引用 3,935 次
- Reducing communication in graph neural network trainingAlok Tripathy, Katherine A. Yelick, Aydin BuluçSC 2020 · 被引用 67 次
- FeatGraph: a flexible and efficient backend for graph neural network systemsYuwei Hu, Zihao Ye, Minjie Wang, Jiali Yu 等SC 2020 · 被引用 57 次
相关 Paper
- DGCL: an efficient communication library for distributed GNN trainingZhenkun Cai, Xiao Yan, Yidi Wu, Kaihao Ma 等EuroSys 2021 · 被引用 103 次
- WholeGraph: A Fast Graph Neural Network Training Framework with Multi-GPU Distributed Shared Memory ArchitectureDongxu Yang, Junhong Liu, Jiaxing Qi, Junjie LaiSC 2022 · 被引用 12 次
- Efficient scaling of dynamic graph neural networksVenkatesan T. Chakaravarthy, Shivmaran S. Pandian, Saurabh Raje, Yogish Sabharwal 等SC 2021 · 被引用 35 次
- ParGNN: A Scalable Graph Neural Network Training Framework on multi-GPUsJunyu Gu, Shunde Li, Rongqiang Cao, Jue Wang 等DAC 2025
- Optimizing Task Placement and Online Scheduling for Distributed GNN Training AccelerationZiyue Luo, Yixin Bao, Chuan WuINFOCOM 2022 · 被引用 12 次
