DynaHB: A Communication-Avoiding Asynchronous Distributed Framework with Hybrid Batches for Dynamic GNN Training
Zhen Song, Yu Gu, Qing Sun, Tianyi Li, Yanfeng Zhang, Yushuai Li, Christian S. Jensen, Ge Yu
摘要
Dynamic Graph Neural Networks (DGNNs) have demonstrated exceptional performance at dynamic-graph analysis tasks. However, the costs exceed those incurred by other learning tasks, to the point where deployment on large-scale dynamic graphs is infeasible. Existing distributed frameworks that facilitate DGNN training are in their early stages and experience challenges such as communication bottlenecks, imbalanced workloads, and GPU memory overflow. We introduce DynaHB, a distributed framework for DGNN training using so-called Hybrid Batches. DynaHB reduces communication by means of vertex caching, and it ensures even data and workload distribution by means of load-aware vertex partitioning. DyanHB also features a novel hybrid-batch training mode that combines vertex-batch and snapshot-batch techniques, thereby reducing training time and GPU memory usage. Next, to further enhance the hybrid batch based approach, DynaHB integrates a reinforcement learning-based batch adjuster and a pipelined batch generator with a batch reservoir to reduce the cost of generating hybrid batches. Extensive experiments show that DynaHB is capable of up to a 93× and an average of 8.06× speedups over the state-of-the-art training framework.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Faster Convergence in Mini-batch Graph Neural Networks Training with Pseudo Full Neighborhood CompensationQiqi Zhou, Yanyan Shen, Lei ChenVLDB 2025 · 被引用 2 次
- Understanding Evolving Graph Structures for Large Discrete-Time Dynamic Graph RepresentationDanni Wu, Yuanyuan Xu, Xuemin Lin, Wenjie Zhang 等VLDB 2026
- PRISM: A Training System to Unlock the Potential of Temporal Graph Learning Through Staleness AvoidanceMd Ashraful Islam, Hojae Son, Suhaas Kiran Doddagaddavalli Gangadharaiah, Marco SerafiniVLDB 2026
- Multimodal Knowledge Graph Completion via Relation-Aware Negative Sampling with Diffusion-based InterpolationQian Ma, Linfei Dai, Zhongming Yao, Yu Gu 等VLDB 2026
- FlareDTDG: Harnessing Temporal Recency for Scalable Discrete-Time Dynamic Graph TrainingWenjie Huang, Rui Wang, Jing Cao, Tongya Zheng 等VLDB 2026
它引用的顶会 Paper30
- GMAN: A Graph Multi-Attention Network for Traffic PredictionChuanpan Zheng, Xiaoliang Fan, Cheng Wang, Jianzhong QiAAAI 2020 · 被引用 1,858 次
- EvolveGCN: Evolving Graph Convolutional Networks for Dynamic GraphsAldo Pareja, Giacomo Domeniconi, Jie Chen, Tengfei Ma 等AAAI 2020 · 被引用 1,429 次
- Inductive representation learning on temporal graphsDa Xu, Chuanwei Ruan, Evren Körpeoglu, Sushant Kumar 等ICLR 2020 · 被引用 901 次
- Inductive Representation Learning in Temporal Networks via Causal Anonymous WalksYanbang Wang, Yen-Yu Chang, Yunyu Liu, Jure Leskovec 等ICLR 2021 · 被引用 326 次
- Towards Better Dynamic Graph Learning: New Architecture and Unified LibraryLe Yu, Leilei Sun, Bowen Du, Weifeng LvNeurIPS 2023 · 被引用 323 次
相关 Paper
- SWASH: A Flexible Communication Framework with Sliding Window-Based Cache Sharing for Scalable DGNN TrainingZhen Song, Yu Gu, Tianyi Li, Yushuai Li 等SIGMOD 2025 · 被引用 3 次
- Scaling New Heights: Transformative Cross-GPU Sampling for Training Billion-Edge GraphsYaqi Xia, Donglin Yang, Xiaobo Zhou, Dazhao ChengSC 2024 · 被引用 4 次
- Efficient scaling of dynamic graph neural networksVenkatesan T. Chakaravarthy, Shivmaran S. Pandian, Saurabh Raje, Yogish Sabharwal 等SC 2021 · 被引用 35 次
- DGC: Training Dynamic Graphs with Spatio-Temporal Non-Uniformity using Graph Partitioning by ChunksFahao Chen, Peng Li, Celimuge WuSIGMOD 2024 · 被引用 10 次
- Two-level Graph Caching for Expediting Distributed GNN TrainingZhe Zhang, Ziyue Luo, Chuan WuINFOCOM 2023 · 被引用 9 次
