BLAD: Adaptive Load Balanced Scheduling and Operator Overlap Pipeline For Accelerating The Dynamic GNN Training
Kaihua Fu, Quan Chen, Yuzhuo Yang, Jiuchen Shi, Chao Li, Minyi Guo
摘要
Dynamic graph networks are widely used for learning time-evolving graphs, but prior work on training these networks is inefficient due to communication overhead, long synchronization, and poor resource usage. Our investigation shows that communication and synchronization can be reduced by carefully scheduling the workload. And the execution order of operators in GNNs can be adjusted without hurting training convergence. We propose a system called BLAD to consider the above factors, comprising a two-level load scheduler and an overlap-aware topology manager. The scheduler allocates each snapshot group to a GPU, alleviating cross-GPU communication. The snapshots in a group are then carefully allocated to processes on a GPU, enabling overlap of compute-intensive NN operators and memory-intensive graph operators. The topology manager adjusts the operators' execution order to maximize the overlap. Experiments show that BLAD achieves 27.2% speed up on training time on average without affecting final accuracy, compared to state-of-the-art solutions.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper5
- Buffalo: Enabling Large-Scale GNN Training via Memory-Efficient BucketizationShuangyan Yang, Minjia Zhang, Dong LiHPCA 2025 · 被引用 10 次
- DynaHB: A Communication-Avoiding Asynchronous Distributed Framework with Hybrid Batches for Dynamic GNN TrainingZhen Song, Yu Gu, Qing Sun, Tianyi Li 等VLDB 2024 · 被引用 7 次
- PRISM: A Training System to Unlock the Potential of Temporal Graph Learning Through Staleness AvoidanceMd Ashraful Islam, Hojae Son, Suhaas Kiran Doddagaddavalli Gangadharaiah, Marco SerafiniVLDB 2026
- PipeTGL: (Near) Zero Bubble Memory-based Temporal Graph Neural Network Training via Pipeline OptimizationJun Liu, Bingqian Du, Ziyue Luo, Sitian Lu 等VLDB 2025
- FlareDTDG: Harnessing Temporal Recency for Scalable Discrete-Time Dynamic Graph TrainingWenjie Huang, Rui Wang, Jing Cao, Tongya Zheng 等VLDB 2026
相关 Paper
- Redundancy-Free High-Performance Dynamic GNN Training with Hierarchical Pipeline ParallelismYaqi Xia, Zheng Zhang, Hulin Wang, Donglin Yang 等HPDC 2023 · 被引用 17 次
- DistTGL: Distributed Memory-Based Temporal Graph Neural Network TrainingHongkuan Zhou, Da Zheng, Xiang Song, George Karypis 等SC 2023 · 被引用 21 次
- Efficient scaling of dynamic graph neural networksVenkatesan T. Chakaravarthy, Shivmaran S. Pandian, Saurabh Raje, Yogish Sabharwal 等SC 2021 · 被引用 35 次
- MGG: Accelerating Graph Neural Networks with Fine-Grained Intra-Kernel Communication-Computation Pipelining on Multi-GPU PlatformsYuke Wang, Boyuan Feng, Zheng Wang, Tong Geng 等OSDI 2023 · 被引用 46 次
- SIMPLE: Efficient Temporal Graph Neural Network Training at Scale with Dynamic Data PlacementShihong Gao, Yiming Li, Xin Zhang, Yanyan Shen 等SIGMOD 2024 · 被引用 19 次
