VeloxGNN: Efficient Out-of-Core GNN Training with Delayed Gradient Propagation
Yi Li, Tsun-Yu Yang, Zhaoyan Shen, Ming-Chang Yang, Bingzhe Li
Abstract
Training Graph Neural Networks (GNNs) on large-scale data is essential in various applications, e.g., transportation, and molecular biology. As graph sizes increasingly surpass main memory capacities, the out-of-core (OOC) GNN training system (OOC-based) has been proposed, a scheme that storing graphs on external storage, such as SSD or HDD, sequentially loading and processing smaller partitions. However, existing OOC-based systems face the key challenges: excessive data migration between storage and memory, as well as reduced model accuracy. In this paper, we theoretically and empirically analyze the limitations of state-of-the-art OOC-based systems and identify opportunities for optimization. Guided by our theoretical insights, we propose VeloxGNN, a novel system to improve data migration efficiency while maintaining high model accuracy. First, we introduce a novel algorithm, named Delayed Gradient Propagation (DGP), which specifically designed for OOC-based system. DGP leverages both historical node embeddings and unbiased gradients to achieve two key objectives simultaneously: minimizing the data migration (by ensuring that dataset is read at most once) while maintaining model accuracy. To support DGP, we then propose system-level optimizations: dynamic memory management, DGP-aware loading order, and a new graph partitioning method that separates labeled and unlabeled data. Experimental results show that VeloxGNN achieves memory-based accuracy while reducing training time by 17.7% to 73.3% across various datasets and GNN models, outperforming state-of-the-art methods. This result highlights VeloxGNN's potential for efficient and scalable GNN training on large-scale graph data.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 2e3af31a-296b-4f4e-81f6-bfe7ee4d3171Related papers
- Celeritas: Out-of-Core Based Unsupervised Graph Neural Network via Cross-Layer Computing 2024Yi Li, Tsun-Yu Yang, Ming-Chang Yang, Zhaoyan Shen et al.HPCA 2024 · 6 citations
- DiskGNN: Bridging I/O Efficiency and Model Accuracy for Out-of-Core GNN TrainingRenjie Liu, Yichuan Wang, Xiao Yan, Haitian Jiang et al.SIGMOD 2025 · 8 citations
- Optimizing Task Placement and Online Scheduling for Distributed GNN Training AccelerationZiyue Luo, Yixin Bao, Chuan WuINFOCOM 2022 · 12 citations
- FreshGNN: Reducing Memory Access via Stable Historical Embeddings for Graph Neural Network TrainingKezhao Huang, Haitian Jiang, Minjie Wang, Guangxuan Xiao et al.VLDB 2024 · 13 citations
- Accelerating Storage-based Training for Graph Neural NetworksMyung-Hwan Jang, Jeong-Min Park, Yunyong Ko, Sang-Wook KimKDD 2026
