Mithril: A Scalable System for Deep GNN Training
Jingji Chen, Zhuoming Chen, Xuehai Qian
Abstract
Communication is a key bottleneck for distributed graph neural network (GNN) training. Existing GNN training systems fail to scale to deep GNNs because of the tremendous amount of inter-GPU communication. This paper proposes Mithril, a new approach that significantly scales the distributed full-graph deep GNN training. Being the first to use layer-level model parallelism for GNN training, Mithril partitions GNN layers among GPUs, each device performs the computation for a disjoint subset of consecutive GNN layers on the whole graph. Compared to graph parallelism with each GPU handling a graph partition, Mithril reduces the communication volume by a factor of the number of GNN layers to scale to deep models. Mithril overcomes the unique challenges for pipelined layer-level model parallelism on the whole graph by partitioning it into dependent chunks, breaking the dependencies with embedding speculation, and applying specific training techniques to ensure convergence. We also propose a hybrid approach by combining Mithril with graph parallelism to handle large graphs, achieve better computer resource utilization and ensure model convergence. We build a general GNN training system supporting all three parallelism settings. Extensive experiments show that Mithril reduces the perepoch communication volume by up to (on average ). It achieves a maximum training time speedup of (on average ) on a GPU cluster with a high-performance InfiniBand network. On another cluster with a commodity Ethernet, Mithril outperforms the baseline by up to (on average ). Mithril also achieves a comparable level of model accuracy and convergence speed compared to graph parallelism.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get c6b8cea3-141f-485f-b94d-e1975d6701ceRelated papers
- Scalable and Efficient Full-Graph GNN Training for Large GraphsXinchen Wan, Kaiqiang Xu, Xudong Liao, Yilun Jin et al.SIGMOD 2023 · 52 citations
- Plexus: Taming Billion-edge Graphs with 3D Parallel Full-graph GNN TrainingAditya K. Ranjan, Siddharth Singh, Cunyang Wei, Abhinav BhateleSC 2025 · 1 citation
- ParGNN: A Scalable Graph Neural Network Training Framework on multi-GPUsJunyu Gu, Shunde Li, Rongqiang Cao, Jue Wang et al.DAC 2025
- Reducing communication in graph neural network trainingAlok Tripathy, Katherine A. Yelick, Aydin BuluçSC 2020 · 67 citations
- Scalable Graph Convolutional Network Training on Distributed-Memory SystemsGunduz Vehbi Demirci, Aparajita Haldar, Hakan FerhatosmanogluVLDB 2023 · 18 citations
