SC2024Top-tier venue
Scaling New Heights: Transformative Cross-GPU Sampling for Training Billion-Edge Graphs
Yaqi Xia, Donglin Yang, Xiaobo Zhou, Dazhao Cheng
Abstract
Efficient training of Graph Neural Networks (GNNs) on billion-edge graphs poses significant challenges due to memory constraints and data transfer bottlenecks, particularly affecting GPU-based sampling. Traditional methods either face severe CPU-GPU data transfer bottlenecks or encounter excessive data shuffling and synchronization overheads in multi-GPU setups. To overcome these challenges in GNN training on large-scale graphs, we introduce HyDRA, a pioneering framework that elevates mini-batch, sampling-based training. HyDRA innovates in multi-GPU memory sharing and multi-node feature retrieval, transforming cross-GPU sampling by seamlessly integrating sampling and data transfer into a single kernel operation. It develops a hybrid pointer-driven data placement technique to enhance neighbor retrieval efficiency, designs a targeted replication strategy for high-degree vertices to reduce communication overhead, and leverages dynamic cross-batch data orchestration with pipelining to minimize redundant data transfers. Evaluated on systems equipped with up to 64 A100 GPUs, HyDRA significantly outperforms current leading methods, achieving to 5.3x faster training speeds compared to DSP and DGL-UVA and demonstrating up to a 42x improvement in multi-GPU scalability. HyDRA sets a new benchmark for high-performance GNN training at large scales.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 987e65a0-8272-4783-ad84-a3c0072eef67Cited by top-tier papers4
- Helios: Efficient Distributed Dynamic Graph Sampling for Online GNN InferenceJie Sun, Zuocheng Shi, Li Su, Wenting Shen et al.PPoPP 2025 · 13 citations
- Voltrix: Sparse Matrix-Matrix Multiplication on Tensor Cores with Asynchronous and Balanced Kernel OptimizationYaqi Xia, Weihu Wang, Donglin Yang, Xiaobo Zhou et al.USENIX ATC 2025 · 6 citations
- Optimizing Data Distribution and Kernel Performance for Efficient Training of Chemistry Foundation Models: A Case Study with MACEJesun Sahariar Firoz, Franco Pellegrini, Mario Geiger, Darren Hsu et al.HPDC 2025 · 2 citations
- Effective and Efficient Distributed Temporal Graph Learning through Hotspot Memory SharingLongjiao Zhang, Rui Wang, Tongya Zheng, Ziqi Huang et al.VLDB 2025 · 1 citation
Related papers
- FastGL: A GPU-Efficient Framework for Accelerating Sampling-Based GNN Training at Large ScaleZeyu Zhu, Peisong Wang, Qinghao Hu, Gang Li et al.ASPLOS 2024 · 8 citations
- Efficient scaling of dynamic graph neural networksVenkatesan T. Chakaravarthy, Shivmaran S. Pandian, Saurabh Raje, Yogish Sabharwal et al.SC 2021 · 35 citations
- DSP: Efficient GNN Training with Multiple GPUsZhenkun Cai, Qihui Zhou, Xiao Yan, Da Zheng et al.PPoPP 2023 · 33 citations
- Global Neighbor Sampling for Mixed CPU-GPU Training on Giant GraphsJialin Dong, Da Zheng, Lin F. Yang, George KarypisKDD 2021 · 22 citations
- Accelerating Sampling and Aggregation Operations in GNN Frameworks with GPU Initiated Direct Storage AccessesJeongmin Brian Park, Vikram Sharma Mailthody, Zaid Qureshi, Wen-Mei HwuVLDB 2024 · 37 citations
