Plexus: Taming Billion-edge Graphs with 3D Parallel Full-graph GNN Training
Aditya K. Ranjan, Siddharth Singh, Cunyang Wei, Abhinav Bhatele
摘要
Graph neural networks (GNNs) leverage the connectivity and structure of real-world graphs to learn intricate properties and relationships between nodes. Many real-world graphs exceed the memory capacity of a GPU due to their sheer size, and training GNNs on such graphs requires techniques such as mini-batch sampling to scale. The alternative approach of distributed full-graph training suffers from high communication overheads and load imbalance due to the irregular structure of graphs. We propose a three-dimensional (3D) parallel approach for full-graph training that tackles these issues and scales to billion-edge graphs. In addition, we introduce optimizations such as a double permutation scheme for load balancing, and a performance model to predict the optimal 3D configuration of our parallel implementation – Plexus. We evaluate Plexus on six different graph datasets and show scaling results on up to 2048 GPUs of Perlmutter, and 1024 GPUs of Frontier. Plexus achieves unprecedented speedups of 2.3 − 12.5 × over prior state of the art, and a reduction in time-to-solution by 5.2 − 8.7 × on Perlmutter and 7.0 − 54.2 × on Frontier.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong 等NeurIPS 2020 · 被引用 3,935 次
- DGCL: an efficient communication library for distributed GNN trainingZhenkun Cai, Xiao Yan, Yidi Wu, Kaihao Ma 等EuroSys 2021 · 被引用 103 次
- PipeGCN: Efficient Full-Graph Training of Graph Convolutional Networks with Pipelined Feature CommunicationCheng Wan, Youjie Li, Cameron R. Wolfe, Anastasios Kyrillidis 等ICLR 2022 · 被引用 89 次
- SANCUS: Staleness-Aware Communication-Avoiding Full-Graph Decentralized Training in Large-Scale Graph Neural NetworksJingshu Peng, Zhao Chen, Yingxia Shao, Yanyan Shen 等VLDB 2022 · 被引用 76 次
- Alpa: Automating Inter- and Intra-Operator Parallelism for Distributed Deep LearningLianmin Zheng, Zhuohan Li, Hao Zhang, Yonghao Zhuang 等OSDI 2022 · 被引用 75 次
相关 Paper
- Scalable and Efficient Full-Graph GNN Training for Large GraphsXinchen Wan, Kaiqiang Xu, Xudong Liao, Yilun Jin 等SIGMOD 2023 · 被引用 52 次
- P3: Distributed Deep Graph Learning at ScaleSwapnil Gandhi, Anand Padmanabha IyerOSDI 2021 · 被引用 192 次
- Mithril: A Scalable System for Deep GNN TrainingJingji Chen, Zhuoming Chen, Xuehai QianHPCA 2025 · 被引用 1 次
- Reducing communication in graph neural network trainingAlok Tripathy, Katherine A. Yelick, Aydin BuluçSC 2020 · 被引用 67 次
- HongTu: Scalable Full-Graph GNN Training on Multiple GPUsQiange Wang, Yao Chen, Weng-Fai Wong, Bingsheng HeSIGMOD 2024 · 被引用 24 次
