SC2025Top-tier venue
Plexus: Taming Billion-edge Graphs with 3D Parallel Full-graph GNN Training
Aditya K. Ranjan, Siddharth Singh, Cunyang Wei, Abhinav Bhatele
Abstract
Graph neural networks (GNNs) leverage the connectivity and structure of real-world graphs to learn intricate properties and relationships between nodes. Many real-world graphs exceed the memory capacity of a GPU due to their sheer size, and training GNNs on such graphs requires techniques such as mini-batch sampling to scale. The alternative approach of distributed full-graph training suffers from high communication overheads and load imbalance due to the irregular structure of graphs. We propose a three-dimensional (3D) parallel approach for full-graph training that tackles these issues and scales to billion-edge graphs. In addition, we introduce optimizations such as a double permutation scheme for load balancing, and a performance model to predict the optimal 3D configuration of our parallel implementation – Plexus. We evaluate Plexus on six different graph datasets and show scaling results on up to 2048 GPUs of Perlmutter, and 1024 GPUs of Frontier. Plexus achieves unprecedented speedups of 2.3 − 12.5 × over prior state of the art, and a reduction in time-to-solution by 5.2 − 8.7 × on Perlmutter and 7.0 − 54.2 × on Frontier.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bc09e867-8932-4f66-a278-b9eec1e9098bBuilds on11
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong et al.NeurIPS 2020 · 3,935 citations
- DGCL: an efficient communication library for distributed GNN trainingZhenkun Cai, Xiao Yan, Yidi Wu, Kaihao Ma et al.EuroSys 2021 · 103 citations
- PipeGCN: Efficient Full-Graph Training of Graph Convolutional Networks with Pipelined Feature CommunicationCheng Wan, Youjie Li, Cameron R. Wolfe, Anastasios Kyrillidis et al.ICLR 2022 · 89 citations
- SANCUS: Staleness-Aware Communication-Avoiding Full-Graph Decentralized Training in Large-Scale Graph Neural NetworksJingshu Peng, Zhao Chen, Yingxia Shao, Yanyan Shen et al.VLDB 2022 · 76 citations
- Alpa: Automating Inter- and Intra-Operator Parallelism for Distributed Deep LearningLianmin Zheng, Zhuohan Li, Hao Zhang, Yonghao Zhuang et al.OSDI 2022 · 75 citations
Related papers
- Scalable and Efficient Full-Graph GNN Training for Large GraphsXinchen Wan, Kaiqiang Xu, Xudong Liao, Yilun Jin et al.SIGMOD 2023 · 52 citations
- P3: Distributed Deep Graph Learning at ScaleSwapnil Gandhi, Anand Padmanabha IyerOSDI 2021 · 192 citations
- Mithril: A Scalable System for Deep GNN TrainingJingji Chen, Zhuoming Chen, Xuehai QianHPCA 2025 · 1 citation
- Reducing communication in graph neural network trainingAlok Tripathy, Katherine A. Yelick, Aydin BuluçSC 2020 · 67 citations
- HongTu: Scalable Full-Graph GNN Training on Multiple GPUsQiange Wang, Yao Chen, Weng-Fai Wong, Bingsheng HeSIGMOD 2024 · 24 citations
