Accelerating Storage-based Training for Graph Neural Networks
Myung-Hwan Jang, Jeong-Min Park, Yunyong Ko, Sang-Wook Kim
Abstract
Graph neural networks (GNNs) have achieved breakthroughs in various real-world downstream tasks due to their powerful expressiveness. As the scale of real-world graphs has been continuously growing, a storage-based approach to GNN training has been studied, which leverages external storage (e.g., NVMe SSDs) to handle such web-scale graphs on a single machine. Although such storagebased GNN training methods have shown promising potential in large-scale GNN training, we observed that they suffer from a severe bottleneck in data preparation since they overlook a critical challenge: how to handle a large number of small storage I/Os. To address the challenge, in this paper, we propose a novel storage-based GNN training framework, named AGNES, that employs a method of block-wise storage I/O processing to fully utilize the I/O bandwidth of high-performance storage devices. Moreover, to further enhance the efficiency of each storage I/O, AGNES employs a simple yet effective strategy, hyperbatch-based processing based on the characteristics of real-world graphs. Comprehensive experiments on five real-world graphs reveal that AGNES consistently outperforms four state-of-the-art methods, up to 4.1× faster than the best competitor. CCS Concepts • Information systems → Data management systems; • Computing methodologies → Machine learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e7f4ecd9-62ab-4800-adad-03f50cb6f804Builds on12
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong et al.NeurIPS 2020 · 3,935 citations
- DropEdge: Towards Deep Graph Convolutional Networks on Node ClassificationYu Rong, Wenbing Huang, Tingyang Xu, Junzhou HuangICLR 2020 · 1,599 citations
- Pick and Choose: A GNN-based Imbalanced Learning Approach for Fraud DetectionYang Liu, Xiang Ao, Zidi Qin, Jianfeng Chi et al.WWW 2021 · 527 citations
- Neo-GNNs: Neighborhood Overlap-aware Graph Neural Networks for Link PredictionSeongjun Yun, Seoyoon Kim, Junhyun Lee, Jaewoo Kang et al.NeurIPS 2021 · 183 citations
- PaGE-Link: Path-based Graph Neural Network Explanation for Heterogeneous Link PredictionShichang Zhang, Jiani Zhang, Xiang Song, Soji Adeshina et al.WWW 2023 · 59 citations
Related papers
- SmartSAGE: training large-scale graph neural networks using in-storage processing architecturesYunjae Lee, Jinha Chung, Minsoo RhuISCA 2022 · 57 citations
- FlashGNN: An In-SSD Accelerator for GNN TrainingFuping Niu, Jianhui Yue, Jiangqiu Shen, Xiaofei Liao et al.HPCA 2024 · 13 citations
- Ginex: SSD-enabled Billion-scale Graph Neural Network Training on a Single Machine via Provably Optimal In-memory CachingYeonhong Park, Sunhong Min, Jae W. LeeVLDB 2022 · 57 citations
- Scaling New Heights: Transformative Cross-GPU Sampling for Training Billion-Edge GraphsYaqi Xia, Donglin Yang, Xiaobo Zhou, Dazhao ChengSC 2024 · 4 citations
- CaliEX: A Disk-Based Large-Scale GNN Training System with Joint Design of Caching and ExecutionCan Su, Haipeng Zhang, Hanyu Zhao, Wenting Shen et al.ICDE 2025 · 1 citation
