FreshGNN: Reducing Memory Access via Stable Historical Embeddings for Graph Neural Network Training
Kezhao Huang, Haitian Jiang, Minjie Wang, Guangxuan Xiao, David Wipf, Xiang Song, Quan Gan, Zengfeng Huang, Jidong Zhai, Zheng Zhang
Abstract
A key performance bottleneck when training graph neural network (GNN) models on large, real-world graphs is loading node features onto a GPU. Due to limited GPU memory, expensive data movement is necessary to facilitate the storage of these features on alternative devices with slower access (e.g. CPU memory). Moreover, the irregularity of graph structures contributes to poor data locality which further exacerbates the problem. Consequently, existing frameworks capable of efficiently training large GNN models usually incur a significant accuracy degradation because of the currently-available shortcuts involved. To address these limitations, we instead propose FreshGNN, a general-purpose GNN mini-batch training framework that leverages a historical cache for storing and reusing GNN node embeddings instead of re-computing them through fetching raw features at every iteration. Critical to its success, the corresponding cache policy is designed, using a combination of gradient-based and staleness criteria, to selectively screen those embeddings which are relatively stable and can be cached, from those that need to be re-computed to reduce estimation errors and subsequent downstream accuracy loss. When paired with complementary system enhancements to support this selective historical cache, FreshGNN is able to accelerate the training speed on large graph datasets such as ogbn-papers100M and MAG240M by 3.4× up to 20.5× and reduce the memory access by 59%, with less than 1% influence on test accuracy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cd232dd1-6ad8-4399-90fb-99d119d44133Cited by top-tier papers5
- A Comprehensive Benchmark on Spectral GNNs: The Impact on Efficiency, Memory, and EffectivenessNingyi Liao, Haoyu Liu, Zulun Zhu, Siqiang Luo et al.SIGMOD 2026 · 4 citations
- Faster Convergence in Mini-batch Graph Neural Networks Training with Pseudo Full Neighborhood CompensationQiqi Zhou, Yanyan Shen, Lei ChenVLDB 2025 · 2 citations
- HypeReca: Distributed Heterogeneous In-Memory Embedding Database for Training Recommender ModelsJiaao He, Shengqi Chen, Kezhao Huang, Jidong ZhaiUSENIX ATC 2025 · 2 citations
- SNI-GNN: SmartNIC-Assisted Full-Graph GNN Training with In-Network Embedding PredictionGuofan Yu, Sitian Chen, Zhenheng Tang, Xiaowen Chu et al.ICDE 2026
- MuseGNN: Forming Scalable, Convergent GNN Layers that Minimize a Sampling-Based EnergyHaitian Jiang, Renjie Liu, Zengfeng Huang, Yichuan Wang et al.ICLR 2025
Builds on17
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong et al.NeurIPS 2020 · 3,935 citations
- GraphSAINT: Graph Sampling Based Inductive Learning MethodHanqing Zeng, Hongkuan Zhou, Ajitesh Srivastava, Rajgopal Kannan et al.ICLR 2020 · 1,155 citations
- Design Space for Graph Neural NetworksJiaxuan You, Zhitao Ying, Jure LeskovecNeurIPS 2020 · 409 citations
- Decoupling the Depth and Scope of Graph Neural NetworksHanqing Zeng, Muhan Zhang, Yinglong Xia, Ajitesh Srivastava et al.NeurIPS 2021 · 189 citations
- Dorylus: Affordable, Scalable, and Accurate GNN Training with Distributed CPU Servers and Serverless ThreadsJohn Thorpe, Yifan Qiao, Jonathan Eyolfson, Shen Teng et al.OSDI 2021 · 175 citations
Related papers
- Haste Makes Waste: A Simple Approach for Scaling Graph Neural NetworksRui Xue, Tong Zhao, Neil Shah, Xiaorui LiuICML 2025
- Accelerating Sampling and Aggregation Operations in GNN Frameworks with GPU Initiated Direct Storage AccessesJeongmin Brian Park, Vikram Sharma Mailthody, Zaid Qureshi, Wen-Mei HwuVLDB 2024 · 37 citations
- DiskGNN: Bridging I/O Efficiency and Model Accuracy for Out-of-Core GNN TrainingRenjie Liu, Yichuan Wang, Xiao Yan, Haitian Jiang et al.SIGMOD 2025 · 8 citations
- FastGL: A GPU-Efficient Framework for Accelerating Sampling-Based GNN Training at Large ScaleZeyu Zhu, Peisong Wang, Qinghao Hu, Gang Li et al.ASPLOS 2024 · 8 citations
- Two-level Graph Caching for Expediting Distributed GNN TrainingZhe Zhang, Ziyue Luo, Chuan WuINFOCOM 2023 · 9 citations
