Lune

HPCA2026顶会

VeloxGNN: Efficient Out-of-Core GNN Training with Delayed Gradient Propagation

Yi Li, Tsun-Yu Yang, Zhaoyan Shen, Ming-Chang Yang, Bingzhe Li

2026年份

摘要

Training Graph Neural Networks (GNNs) on large-scale data is essential in various applications, e.g., transportation, and molecular biology. As graph sizes increasingly surpass main memory capacities, the out-of-core (OOC) GNN training system (OOC-based) has been proposed, a scheme that storing graphs on external storage, such as SSD or HDD, sequentially loading and processing smaller partitions. However, existing OOC-based systems face the key challenges: excessive data migration between storage and memory, as well as reduced model accuracy. In this paper, we theoretically and empirically analyze the limitations of state-of-the-art OOC-based systems and identify opportunities for optimization. Guided by our theoretical insights, we propose VeloxGNN, a novel system to improve data migration efficiency while maintaining high model accuracy. First, we introduce a novel algorithm, named Delayed Gradient Propagation (DGP), which specifically designed for OOC-based system. DGP leverages both historical node embeddings and unbiased gradients to achieve two key objectives simultaneously: minimizing the data migration (by ensuring that dataset is read at most once) while maintaining model accuracy. To support DGP, we then propose system-level optimizations: dynamic memory management, DGP-aware loading order, and a new graph partitioning method that separates labeled and unlabeled data. Experimental results show that VeloxGNN achieves memory-based accuracy while reducing training time by 17.7% to 73.3% across various datasets and GNN models, outperforming state-of-the-art methods. This result highlights VeloxGNN's potential for efficient and scalable GNN training on large-scale graph data.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖