Accelerating Sampling and Aggregation Operations in GNN Frameworks with GPU Initiated Direct Storage Accesses
Jeongmin Brian Park, Vikram Sharma Mailthody, Zaid Qureshi, Wen-Mei Hwu
摘要
Graph Neural Networks (GNNs) are emerging as a powerful tool for learning from graph-structured data and performing sophisticated inference tasks in various application domains. Although GNNs have been shown to be effective on modest-sized graphs, training them on large-scale graphs remains a significant challenge due to the lack of efficient storage access and caching methods for graph data. Existing frameworks for training GNNs use CPUs for graph sampling and feature aggregation, while the training and updating of model weights are executed on GPUs. However, our in-depth profiling shows CPUs cannot achieve the graph sampling and feature aggregation throughput required to keep up with GPUs. Furthermore, when the graph and its embeddings do not fit in the CPU memory, the overhead introduced by the operating system, say for handling page-faults, causes gross under-utilization of hardware and prolonged end-to-end execution time. To address these issues, we propose the GPU Initiated Direct Storage Access (GIDS) dataloader, to enable GPU-oriented GNN training for large-scale graphs while efficiently utilizing all hardware resources, such as CPU memory, storage, and GPU memory. The GIDS dataloader first addresses memory capacity constraints by enabling GPU threads to directly fetch feature vectors from storage. Then, we introduce a set of innovative solutions, including the dynamic storage access accumulator, constant CPU buffer, and GPU software cache with window buffering, to balance resource utilization across the entire system for improved end-to-end training throughput. Our evaluation using a single GPU on terabyte-scale GNN datasets shows that the GIDS dataloader accelerates the overall DGL GNN training pipeline by up to 582× when compared to the current, state-of-the-art DGL dataloader.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- OUTRE: An OUT-of-core De-REdundancy GNN Training Framework for Massive Graphs within A Single MachineZeang Sheng, Wentao Zhang, Yangyu Tao, Bin CuiVLDB 2024 · 被引用 17 次
- DiskGNN: Bridging I/O Efficiency and Model Accuracy for Out-of-Core GNN TrainingRenjie Liu, Yichuan Wang, Xiao Yan, Haitian Jiang 等SIGMOD 2025 · 被引用 8 次
- Ratel: Optimizing Holistic Data Movement to Fine-tune 100B Model on a Consumer GPUChangyue Liao, Mo Sun, Zihan Yang, Jun Xie 等ICDE 2025 · 被引用 4 次
- CAM: Asynchronous GPU-Initiated, CPU-Managed SSD Management for Batching Storage AccessZiyu Song, Jie Zhang, Jie Sun, Mo Sun 等ICDE 2025 · 被引用 4 次
- LongSight: Compute-Enabled Memory to Accelerate Large-Context LLMs via Sparse AttentionDerrick Quinn, E. Ezgi Yücel, Jinkwon Kim, José F. Martínez 等MICRO 2025 · 被引用 3 次
它引用的顶会 Paper17
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong 等NeurIPS 2020 · 被引用 3,935 次
- ZeRO-infinity: breaking the GPU memory wall for extreme scale deep learningSamyam Rajbhandari, Olatunji Ruwase, Jeff Rasley, Shaden Smith 等SC 2021 · 被引用 254 次
- NeutronStar: Distributed GNN Training with Hybrid Dependency ManagementQiange Wang, Yanfeng Zhang, Hao Wang, Chaoyi Chen 等SIGMOD 2022 · 被引用 60 次
- FeatGraph: a flexible and efficient backend for graph neural network systemsYuwei Hu, Zihao Ye, Minjie Wang, Jiali Yu 等SC 2020 · 被引用 57 次
- Ginex: SSD-enabled Billion-scale Graph Neural Network Training on a Single Machine via Provably Optimal In-memory CachingYeonhong Park, Sunhong Min, Jae W. LeeVLDB 2022 · 被引用 57 次
相关 Paper
- FastGL: A GPU-Efficient Framework for Accelerating Sampling-Based GNN Training at Large ScaleZeyu Zhu, Peisong Wang, Qinghao Hu, Gang Li 等ASPLOS 2024 · 被引用 8 次
- BGL: GPU-Efficient GNN Training by Optimizing Graph Data I/O and PreprocessingTianfeng Liu, Yangrui Chen, Dan Li, Chuan Wu 等NSDI 2023
- Scaling New Heights: Transformative Cross-GPU Sampling for Training Billion-Edge GraphsYaqi Xia, Donglin Yang, Xiaobo Zhou, Dazhao ChengSC 2024 · 被引用 4 次
- FreshGNN: Reducing Memory Access via Stable Historical Embeddings for Graph Neural Network TrainingKezhao Huang, Haitian Jiang, Minjie Wang, Guangxuan Xiao 等VLDB 2024 · 被引用 13 次
- WholeGraph: A Fast Graph Neural Network Training Framework with Multi-GPU Distributed Shared Memory ArchitectureDongxu Yang, Junhong Liu, Jiaxing Qi, Junjie LaiSC 2022 · 被引用 12 次
