Mithril: A Scalable System for Deep GNN Training
Jingji Chen, Zhuoming Chen, Xuehai Qian
摘要
Communication is a key bottleneck for distributed graph neural network (GNN) training. Existing GNN training systems fail to scale to deep GNNs because of the tremendous amount of inter-GPU communication. This paper proposes Mithril, a new approach that significantly scales the distributed full-graph deep GNN training. Being the first to use layer-level model parallelism for GNN training, Mithril partitions GNN layers among GPUs, each device performs the computation for a disjoint subset of consecutive GNN layers on the whole graph. Compared to graph parallelism with each GPU handling a graph partition, Mithril reduces the communication volume by a factor of the number of GNN layers to scale to deep models. Mithril overcomes the unique challenges for pipelined layer-level model parallelism on the whole graph by partitioning it into dependent chunks, breaking the dependencies with embedding speculation, and applying specific training techniques to ensure convergence. We also propose a hybrid approach by combining Mithril with graph parallelism to handle large graphs, achieve better computer resource utilization and ensure model convergence. We build a general GNN training system supporting all three parallelism settings. Extensive experiments show that Mithril reduces the perepoch communication volume by up to (on average ). It achieves a maximum training time speedup of (on average ) on a GPU cluster with a high-performance InfiniBand network. On another cluster with a commodity Ethernet, Mithril outperforms the baseline by up to (on average ). Mithril also achieves a comparable level of model accuracy and convergence speed compared to graph parallelism.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Scalable and Efficient Full-Graph GNN Training for Large GraphsXinchen Wan, Kaiqiang Xu, Xudong Liao, Yilun Jin 等SIGMOD 2023 · 被引用 52 次
- Plexus: Taming Billion-edge Graphs with 3D Parallel Full-graph GNN TrainingAditya K. Ranjan, Siddharth Singh, Cunyang Wei, Abhinav BhateleSC 2025 · 被引用 1 次
- ParGNN: A Scalable Graph Neural Network Training Framework on multi-GPUsJunyu Gu, Shunde Li, Rongqiang Cao, Jue Wang 等DAC 2025
- Reducing communication in graph neural network trainingAlok Tripathy, Katherine A. Yelick, Aydin BuluçSC 2020 · 被引用 67 次
- Scalable Graph Convolutional Network Training on Distributed-Memory SystemsGunduz Vehbi Demirci, Aparajita Haldar, Hakan FerhatosmanogluVLDB 2023 · 被引用 18 次
