Redundancy-Free High-Performance Dynamic GNN Training with Hierarchical Pipeline Parallelism
Yaqi Xia, Zheng Zhang, Hulin Wang, Donglin Yang, Xiaobo Zhou, Dazhao Cheng
Abstract
Temporal Graph Neural Networks(TGNNs) extend the success of Graph Neural Networks to dynamic graphs. Distributed TGNN training requires efficiently tackling temporal dependency, which often leads to excessive cross-device communication that generates significant redundant data. However, existing systems are unable to remove the redundancy in data reuse and transfer, and suffer from severe communication overhead in a distributed setting. This paper presents Sven, an algorithm and system co-designed TGNN training library for the end-to-end performance optimization on multi-node multi-GPU systems. Exploiting dependency patterns of TGNN models and characteristics of dynamic graph datasets, we design redundancy-free data organization and load-balancing partitioning strategies that mitigate the redundant data communication and evenly partition dynamic graphs at the vertex level. Furthermore, we develop a hierarchical pipeline mechanism integrating data prefetching, micro-batch pipelining, and asynchronous pipelining to mitigate the communication overhead. As the first scaling study on the memory-based TGNNs training, experiments conducted on an HPC cluster of 64 GPUs show that Sven can achieve up to 1.7x-3.3x speedup over the state-of-art approaches and a factor of up to 5.26x communication efficiency improvement.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 4dfb87af-05e6-499b-a4f6-b035066d65a5Cited by top-tier papers7
- Helios: Efficient Distributed Dynamic Graph Sampling for Online GNN InferenceJie Sun, Zuocheng Shi, Li Su, Wenting Shen et al.PPoPP 2025 · 13 citations
- Buffalo: Enabling Large-Scale GNN Training via Memory-Efficient BucketizationShuangyan Yang, Minjia Zhang, Dong LiHPCA 2025 · 10 citations
- MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive OperatorsZheng Zhang, Donglin Yang, Xiaobo Zhou, Dazhao ChengSC 2024 · 9 citations
- DynaHB: A Communication-Avoiding Asynchronous Distributed Framework with Hybrid Batches for Dynamic GNN TrainingZhen Song, Yu Gu, Qing Sun, Tianyi Li et al.VLDB 2024 · 7 citations
- Voltrix: Sparse Matrix-Matrix Multiplication on Tensor Cores with Asynchronous and Balanced Kernel OptimizationYaqi Xia, Weihu Wang, Donglin Yang, Xiaobo Zhou et al.USENIX ATC 2025 · 6 citations
Related papers
- Efficient scaling of dynamic graph neural networksVenkatesan T. Chakaravarthy, Shivmaran S. Pandian, Saurabh Raje, Yogish Sabharwal et al.SC 2021 · 35 citations
- DistTGL: Distributed Memory-Based Temporal Graph Neural Network TrainingHongkuan Zhou, Da Zheng, Xiang Song, George Karypis et al.SC 2023 · 21 citations
- PiPAD: Pipelined and Parallel Dynamic GNN Training on GPUsChunyang Wang, Desen Sun, Yuebin BaiPPoPP 2023 · 27 citations
- Efficient Temporal Graph Network Training via Unified Redundancy EliminationYiqing Wang, Hailong Yang, Kejie Ma, Enze Yu et al.ASPLOS 2026 · 1 citation
- PipeTGL: (Near) Zero Bubble Memory-based Temporal Graph Neural Network Training via Pipeline OptimizationJun Liu, Bingqian Du, Ziyue Luo, Sitian Lu et al.VLDB 2025
