Adaptive Parallel Training for Graph Neural Networks
Kaihao Ma, Renjie Liu, Xiao Yan, Zhenkun Cai, Xiang Song, Minjie Wang, Yichao Li, James Cheng
摘要
There are several strategies to parallelize graph neural network (GNN) training over multiple GPUs. We observe that there is no consistent winner (i.e., with the shortest running time), and the optimal strategy depends on the graph dataset, GNN model, training algorithm, and hardware configurations. As such, we design the APT system to automatically select efficient parallelization strategies for GNN training tasks. To this end, we analyze the trade-offs of the strategies and design simple yet effective cost models to compare their execution time and facilitate strategy selection. Moreover, we also propose a general abstraction of the strategies, which allows to implement a unified execution engine that can be configured to run different strategies. Our experiments show that APT usually chooses the optimal or a close to optimal strategy, and the training time can be reduced by over 2x compared with always using a single strategy. APT is open-source at https://github.com/kaihaoma/APT.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- WiseGraph: Optimizing GNN with Joint Workload Partition of Graph and OperationsKezhao Huang, Jidong Zhai, Liyan Zheng, Haojie Wang 等EuroSys 2024 · 被引用 11 次
- CoGNN: Efficient Scheduling for Concurrent GNN Training on GPUsQingxiao Sun, Yi Liu, Hailong Yang, Ruizhe Zhang 等SC 2022 · 被引用 11 次
- ElasGNN: An Elastic Training Framework for Distributed GNN TrainingSiqi Wang, Hailong Yang, Pengbo Wang, Hongliang Cao 等PPoPP 2026
- DGCL: an efficient communication library for distributed GNN trainingZhenkun Cai, Xiao Yan, Yidi Wu, Kaihao Ma 等EuroSys 2021 · 被引用 103 次
- uGrapher: High-Performance Graph Operator Computation via Unified Abstraction for Graph Neural NetworksYangjie Zhou, Jingwen Leng, Yaoxu Song, Shuwen Lu 等ASPLOS 2023 · 被引用 25 次
