Scaling New Heights: Transformative Cross-GPU Sampling for Training Billion-Edge Graphs
Yaqi Xia, Donglin Yang, Xiaobo Zhou, Dazhao Cheng
摘要
Efficient training of Graph Neural Networks (GNNs) on billion-edge graphs poses significant challenges due to memory constraints and data transfer bottlenecks, particularly affecting GPU-based sampling. Traditional methods either face severe CPU-GPU data transfer bottlenecks or encounter excessive data shuffling and synchronization overheads in multi-GPU setups. To overcome these challenges in GNN training on large-scale graphs, we introduce HyDRA, a pioneering framework that elevates mini-batch, sampling-based training. HyDRA innovates in multi-GPU memory sharing and multi-node feature retrieval, transforming cross-GPU sampling by seamlessly integrating sampling and data transfer into a single kernel operation. It develops a hybrid pointer-driven data placement technique to enhance neighbor retrieval efficiency, designs a targeted replication strategy for high-degree vertices to reduce communication overhead, and leverages dynamic cross-batch data orchestration with pipelining to minimize redundant data transfers. Evaluated on systems equipped with up to 64 A100 GPUs, HyDRA significantly outperforms current leading methods, achieving to 5.3x faster training speeds compared to DSP and DGL-UVA and demonstrating up to a 42x improvement in multi-GPU scalability. HyDRA sets a new benchmark for high-performance GNN training at large scales.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper4
- Helios: Efficient Distributed Dynamic Graph Sampling for Online GNN InferenceJie Sun, Zuocheng Shi, Li Su, Wenting Shen 等PPoPP 2025 · 被引用 13 次
- Voltrix: Sparse Matrix-Matrix Multiplication on Tensor Cores with Asynchronous and Balanced Kernel OptimizationYaqi Xia, Weihu Wang, Donglin Yang, Xiaobo Zhou 等USENIX ATC 2025 · 被引用 6 次
- Optimizing Data Distribution and Kernel Performance for Efficient Training of Chemistry Foundation Models: A Case Study with MACEJesun Sahariar Firoz, Franco Pellegrini, Mario Geiger, Darren Hsu 等HPDC 2025 · 被引用 2 次
- Effective and Efficient Distributed Temporal Graph Learning through Hotspot Memory SharingLongjiao Zhang, Rui Wang, Tongya Zheng, Ziqi Huang 等VLDB 2025 · 被引用 1 次
相关 Paper
- FastGL: A GPU-Efficient Framework for Accelerating Sampling-Based GNN Training at Large ScaleZeyu Zhu, Peisong Wang, Qinghao Hu, Gang Li 等ASPLOS 2024 · 被引用 8 次
- Efficient scaling of dynamic graph neural networksVenkatesan T. Chakaravarthy, Shivmaran S. Pandian, Saurabh Raje, Yogish Sabharwal 等SC 2021 · 被引用 35 次
- DSP: Efficient GNN Training with Multiple GPUsZhenkun Cai, Qihui Zhou, Xiao Yan, Da Zheng 等PPoPP 2023 · 被引用 33 次
- Global Neighbor Sampling for Mixed CPU-GPU Training on Giant GraphsJialin Dong, Da Zheng, Lin F. Yang, George KarypisKDD 2021 · 被引用 22 次
- Accelerating Sampling and Aggregation Operations in GNN Frameworks with GPU Initiated Direct Storage AccessesJeongmin Brian Park, Vikram Sharma Mailthody, Zaid Qureshi, Wen-Mei HwuVLDB 2024 · 被引用 37 次
