Ginex: SSD-enabled Billion-scale Graph Neural Network Training on a Single Machine via Provably Optimal In-memory Caching
Yeonhong Park, Sunhong Min, Jae W. Lee
Abstract
Graph Neural Networks (GNNs) are receiving a spotlight as a powerful tool that can effectively serve various inference tasks on graph structured data. As the size of real-world graphs continues to scale, the GNN training system faces a scalability challenge. Distributed training is a popular approach to address this challenge by scaling out CPU nodes. However, not much attention has been paid to disk-based GNN training, which can scale up the single-node system in a more cost-effective manner by leveraging high-performance storage devices like NVMe SSDs. We observe that the data movement between the main memory and the disk is the primary bottleneck in the SSD-based training system, and that the conventional GNN training pipeline is sub-optimal without taking this overhead into account. Thus, we propose Ginex, the first SSD-based GNN training system that can process billion-scale graph datasets on a single machine. Inspired by the inspector-executor execution model in compiler optimization, Ginex restructures the GNN training pipeline by separating sample and gather stages. This separation enables Ginex to realize a provably optimal replacement algorithm, known as Belady's algorithm , for caching feature vectors in memory, which account for the dominant portion of I/O accesses. According to our evaluation with four billion-scale graph datasets and two GNN models, Ginex achieves 2.11X higher training throughput on average (2.67X at maximum) than the SSD-extended PyTorch Geometric.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7cffd3f6-7e86-4b76-89d3-6a9a4563d389Cited by top-tier papers17
- Accelerating Sampling and Aggregation Operations in GNN Frameworks with GPU Initiated Direct Storage AccessesJeongmin Brian Park, Vikram Sharma Mailthody, Zaid Qureshi, Wen-Mei HwuVLDB 2024 · 37 citations
- BeaconGNN: Large-Scale GNN Acceleration with Out-of-Order Streaming In-Storage ComputingYuyue Wang, Xiurui Pan, Yuda An, Jie Zhang et al.HPCA 2024 · 27 citations
- Distributed Graph Embedding with Information-Oriented Random WalksPeng Fang, Arijit Khan, Siqiang Luo, Fang Wang et al.VLDB 2023 · 18 citations
- NeutronStream: A Dynamic GNN Training Framework with Sliding Window for Graph StreamsChaoyi Chen, Dechao Gao, Yanfeng Zhang, Qiange Wang et al.VLDB 2024 · 18 citations
- OUTRE: An OUT-of-core De-REdundancy GNN Training Framework for Massive Graphs within A Single MachineZeang Sheng, Wentao Zhang, Yangyu Tao, Bin CuiVLDB 2024 · 17 citations
Builds on10
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong et al.NeurIPS 2020 · 3,935 citations
- An Imitation Learning Approach for Cache ReplacementEvan Zheran Liu, Milad Hashemi, Kevin Swersky, Parthasarathy Ranganathan et al.ICML 2020 · 108 citations
- Large Graph Convolutional Network Training with GPU-Oriented Data Communication ArchitectureSeungwon Min, Kun Wu, Sitao Huang, Mert Hidayetoglu et al.VLDB 2021 · 85 citations
- Marius: Learning Massive Graph Embeddings on a Single MachineJason Mohoney, Roger Waleffe, Henry Xu, Theodoros Rekatsinas et al.OSDI 2021 · 75 citations
- Minimal Variance Sampling with Provable Guarantees for Fast Training of Graph Neural NetworksWeilin Cong, Rana Forsati, Mahmut T. Kandemir, Mehrdad MahdaviKDD 2020 · 73 citations
Related papers
- CaliEX: A Disk-Based Large-Scale GNN Training System with Joint Design of Caching and ExecutionCan Su, Haipeng Zhang, Hanyu Zhao, Wenting Shen et al.ICDE 2025 · 1 citation
- FlashGNN: An In-SSD Accelerator for GNN TrainingFuping Niu, Jianhui Yue, Jiangqiu Shen, Xiaofei Liao et al.HPCA 2024 · 13 citations
- MariusGNN: Resource-Efficient Out-of-Core Training of Graph Neural NetworksRoger Waleffe, Jason Mohoney, Theodoros Rekatsinas, Shivaram VenkataramanEuroSys 2023 · 40 citations
- Hyperion: Co-Optimizing SSD Access and GPU Computation for Cost-Efficient GNN TrainingJie Sun, Mo Sun, Zheng Zhang, Zuocheng Shi et al.ICDE 2025 · 3 citations
- Accelerating Storage-based Training for Graph Neural NetworksMyung-Hwan Jang, Jeong-Min Park, Yunyong Ko, Sang-Wook KimKDD 2026
