Accelerating graph sampling for graph machine learning using GPUs
Abhinav Jangda, Sandeep Polisetty, Arjun Guha, Marco Serafini
Abstract
Representation learning algorithms automatically learn the features of data. Several representation learning algorithms for graph data, such as DeepWalk, node2vec, and Graph-SAGE, sample the graph to produce mini-batches that are suitable for training a DNN. However, sampling time can be a significant fraction of training time, and existing systems do not efficiently parallelize sampling.
Sampling is an "embarrassingly parallel" problem and may appear to lend itself to GPU acceleration, but the irregularity of graphs makes it hard to use GPU resources effectively. This paper presents NextDoor, a system designed to effectively perform graph sampling on GPUs. NextDoor employs a new approach to graph sampling that we call transit-parallelism, which allows load balancing and caching of edges. NextDoor provides end-users with a high-level abstraction for writing a variety of graph sampling algorithms. We implement several graph sampling applications, and show that NextDoor runs them orders of magnitude faster than existing systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 21b2d901-5d6b-4d39-a44b-d2b5d0681b16Cited by top-tier papers26
- ByteGNN: Efficient Graph Neural Network Training at Large ScaleChenguang Zheng, Hongzhi Chen, Yuxuan Cheng, Zhezheng Song et al.VLDB 2022 · 107 citations
- GNNLab: a factored system for sample-based GNN training over GPUsJianbang Yang, Dahai Tang, Xiaoniu Song, Lei Wang et al.EuroSys 2022 · 105 citations
- PaSca: A Graph Neural Architecture Search System under the Scalable ParadigmWentao Zhang, Yu Shen, Zheyu Lin, Yang Li et al.WWW 2022 · 69 citations
- MariusGNN: Resource-Efficient Out-of-Core Training of Graph Neural NetworksRoger Waleffe, Jason Mohoney, Theodoros Rekatsinas, Shivaram VenkataramanEuroSys 2023 · 40 citations
- Towards Training Billion Parameter Graph Neural Networks for Atomic SimulationsAnuroop Sriram, Abhishek Das, Brandon M. Wood, Siddharth Goyal et al.ICLR 2022 · 39 citations
Builds on5
- GraphSAINT: Graph Sampling Based Inductive Learning MethodHanqing Zeng, Hongkuan Zhou, Ajitesh Srivastava, Rajgopal Kannan et al.ICLR 2020 · 1,155 citations
- Subway: minimizing data transfer during out-of-GPU-memory graph processingAmir Hossein Nodehi Sabet, Zhijia Zhao, Rajiv GuptaEuroSys 2020 · 84 citations
- Pangolin: An Efficient and Flexible Graph Mining System on CPU and GPUXuhao Chen, Roshan Dathathri, Gurbinder Gill, Keshav PingaliVLDB 2020 · 81 citations
- Minimal Variance Sampling with Provable Guarantees for Fast Training of Graph Neural NetworksWeilin Cong, Rana Forsati, Mahmut T. Kandemir, Mehrdad MahdaviKDD 2020 · 73 citations
- SympleGraph: distributed graph processing with precise loop-carried dependency guaranteeYouwei Zhuo, Jingji Chen, Qinyi Luo, Yanzhi Wang et al.PLDI 2020 · 13 citations
Related papers
- gSampler: General and Efficient GPU-based Graph Sampling for Graph LearningPing Gong, Renjie Liu, Zunyao Mao, Zhenkun Cai et al.SOSP 2023 · 18 citations
- C-SAW: a framework for graph sampling and random walk on GPUsSantosh Pandey, Lingda Li, Adolfy Hoisie, Xiaoye S. Li et al.SC 2020 · 51 citations
- DSP: Efficient GNN Training with Multiple GPUsZhenkun Cai, Qihui Zhou, Xiao Yan, Da Zheng et al.PPoPP 2023 · 33 citations
- Self-adaptive Graph Traversal on GPUsMo Sha, Yuchen Li, Kian-Lee TanSIGMOD 2021 · 12 citations
- FastGL: A GPU-Efficient Framework for Accelerating Sampling-Based GNN Training at Large ScaleZeyu Zhu, Peisong Wang, Qinghao Hu, Gang Li et al.ASPLOS 2024 · 8 citations
