SC2023Top-tier venue
BLAD: Adaptive Load Balanced Scheduling and Operator Overlap Pipeline For Accelerating The Dynamic GNN Training
Kaihua Fu, Quan Chen, Yuzhuo Yang, Jiuchen Shi, Chao Li, Minyi Guo
Abstract
Dynamic graph networks are widely used for learning time-evolving graphs, but prior work on training these networks is inefficient due to communication overhead, long synchronization, and poor resource usage. Our investigation shows that communication and synchronization can be reduced by carefully scheduling the workload. And the execution order of operators in GNNs can be adjusted without hurting training convergence. We propose a system called BLAD to consider the above factors, comprising a two-level load scheduler and an overlap-aware topology manager. The scheduler allocates each snapshot group to a GPU, alleviating cross-GPU communication. The snapshots in a group are then carefully allocated to processes on a GPU, enabling overlap of compute-intensive NN operators and memory-intensive graph operators. The topology manager adjusts the operators' execution order to maximize the overlap. Experiments show that BLAD achieves 27.2% speed up on training time on average without affecting final accuracy, compared to state-of-the-art solutions.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 8e659e06-4894-42a6-be09-5ddabe049183Cited by top-tier papers5
- Buffalo: Enabling Large-Scale GNN Training via Memory-Efficient BucketizationShuangyan Yang, Minjia Zhang, Dong LiHPCA 2025 · 10 citations
- DynaHB: A Communication-Avoiding Asynchronous Distributed Framework with Hybrid Batches for Dynamic GNN TrainingZhen Song, Yu Gu, Qing Sun, Tianyi Li et al.VLDB 2024 · 7 citations
- PRISM: A Training System to Unlock the Potential of Temporal Graph Learning Through Staleness AvoidanceMd Ashraful Islam, Hojae Son, Suhaas Kiran Doddagaddavalli Gangadharaiah, Marco SerafiniVLDB 2026
- PipeTGL: (Near) Zero Bubble Memory-based Temporal Graph Neural Network Training via Pipeline OptimizationJun Liu, Bingqian Du, Ziyue Luo, Sitian Lu et al.VLDB 2025
- FlareDTDG: Harnessing Temporal Recency for Scalable Discrete-Time Dynamic Graph TrainingWenjie Huang, Rui Wang, Jing Cao, Tongya Zheng et al.VLDB 2026
Related papers
- Redundancy-Free High-Performance Dynamic GNN Training with Hierarchical Pipeline ParallelismYaqi Xia, Zheng Zhang, Hulin Wang, Donglin Yang et al.HPDC 2023 · 17 citations
- DistTGL: Distributed Memory-Based Temporal Graph Neural Network TrainingHongkuan Zhou, Da Zheng, Xiang Song, George Karypis et al.SC 2023 · 21 citations
- Efficient scaling of dynamic graph neural networksVenkatesan T. Chakaravarthy, Shivmaran S. Pandian, Saurabh Raje, Yogish Sabharwal et al.SC 2021 · 35 citations
- MGG: Accelerating Graph Neural Networks with Fine-Grained Intra-Kernel Communication-Computation Pipelining on Multi-GPU PlatformsYuke Wang, Boyuan Feng, Zheng Wang, Tong Geng et al.OSDI 2023 · 46 citations
- SIMPLE: Efficient Temporal Graph Neural Network Training at Scale with Dynamic Data PlacementShihong Gao, Yiming Li, Xin Zhang, Yanyan Shen et al.SIGMOD 2024 · 19 citations
