Redundancy-Free High-Performance Dynamic GNN Training with Hierarchical Pipeline Parallelism
Yaqi Xia, Zheng Zhang, Hulin Wang, Donglin Yang, Xiaobo Zhou, Dazhao Cheng
摘要
Temporal Graph Neural Networks(TGNNs) extend the success of Graph Neural Networks to dynamic graphs. Distributed TGNN training requires efficiently tackling temporal dependency, which often leads to excessive cross-device communication that generates significant redundant data. However, existing systems are unable to remove the redundancy in data reuse and transfer, and suffer from severe communication overhead in a distributed setting. This paper presents Sven, an algorithm and system co-designed TGNN training library for the end-to-end performance optimization on multi-node multi-GPU systems. Exploiting dependency patterns of TGNN models and characteristics of dynamic graph datasets, we design redundancy-free data organization and load-balancing partitioning strategies that mitigate the redundant data communication and evenly partition dynamic graphs at the vertex level. Furthermore, we develop a hierarchical pipeline mechanism integrating data prefetching, micro-batch pipelining, and asynchronous pipelining to mitigate the communication overhead. As the first scaling study on the memory-based TGNNs training, experiments conducted on an HPC cluster of 64 GPUs show that Sven can achieve up to 1.7x-3.3x speedup over the state-of-art approaches and a factor of up to 5.26x communication efficiency improvement.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper7
- Helios: Efficient Distributed Dynamic Graph Sampling for Online GNN InferenceJie Sun, Zuocheng Shi, Li Su, Wenting Shen 等PPoPP 2025 · 被引用 13 次
- Buffalo: Enabling Large-Scale GNN Training via Memory-Efficient BucketizationShuangyan Yang, Minjia Zhang, Dong LiHPCA 2025 · 被引用 10 次
- MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive OperatorsZheng Zhang, Donglin Yang, Xiaobo Zhou, Dazhao ChengSC 2024 · 被引用 9 次
- DynaHB: A Communication-Avoiding Asynchronous Distributed Framework with Hybrid Batches for Dynamic GNN TrainingZhen Song, Yu Gu, Qing Sun, Tianyi Li 等VLDB 2024 · 被引用 7 次
- Voltrix: Sparse Matrix-Matrix Multiplication on Tensor Cores with Asynchronous and Balanced Kernel OptimizationYaqi Xia, Weihu Wang, Donglin Yang, Xiaobo Zhou 等USENIX ATC 2025 · 被引用 6 次
相关 Paper
- Efficient scaling of dynamic graph neural networksVenkatesan T. Chakaravarthy, Shivmaran S. Pandian, Saurabh Raje, Yogish Sabharwal 等SC 2021 · 被引用 35 次
- DistTGL: Distributed Memory-Based Temporal Graph Neural Network TrainingHongkuan Zhou, Da Zheng, Xiang Song, George Karypis 等SC 2023 · 被引用 21 次
- PiPAD: Pipelined and Parallel Dynamic GNN Training on GPUsChunyang Wang, Desen Sun, Yuebin BaiPPoPP 2023 · 被引用 27 次
- Efficient Temporal Graph Network Training via Unified Redundancy EliminationYiqing Wang, Hailong Yang, Kejie Ma, Enze Yu 等ASPLOS 2026 · 被引用 1 次
- PipeTGL: (Near) Zero Bubble Memory-based Temporal Graph Neural Network Training via Pipeline OptimizationJun Liu, Bingqian Du, Ziyue Luo, Sitian Lu 等VLDB 2025
