PRES: Toward Scalable Memory-Based Dynamic Graph Neural Networks
Junwei Su, Difan Zou, Chuan Wu
Abstract
Memory-based Dynamic Graph Neural Networks (MDGNNs) are a family of dynamic graph neural networks that leverage a memory module to extract, distill, and memorize long-term temporal dependencies, leading to superior performance compared to memory-less counterparts. However, training MDGNNs faces the challenge of handling entangled temporal and structural dependencies, requiring sequential and chronological processing of data sequences to capture accurate temporal patterns. During the batch training, the temporal data points within the same batch will be processed in parallel, while their temporal dependencies are neglected. This issue is referred to as temporal discontinuity and restricts the effective temporal batch size, limiting data parallelism and reducing MDGNNs' flexibility in industrial applications. This paper studies the efficient training of MDGNNs at scale, focusing on the temporal discontinuity in training MDGNNs with large temporal batch sizes. We first conduct a theoretical study on the impact of temporal batch size on the convergence of MDGNN training. Based on the analysis, we propose PRES, an iterative prediction-correction scheme combined with a memory coherence learning objective to mitigate the effect of temporal discontinuity, enabling MDGNNs to be trained with significantly larger temporal batches without sacrificing generalization performance. Experimental results demonstrate that our approach enables up to a 4 × larger temporal batch (3.4× speed-up) during MDGNN training.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers13
- Towards Robust Graph Incremental Learning on Evolving GraphsJunwei Su, Difan Zou, Zijun Zhang, Chuan WuICML 2023 · 37 citations
- MSPipe: Efficient Temporal GNN Training via Staleness-Aware PipelineGuangming Sheng, Junwei Su, Chao Huang, Chuan WuKDD 2024 · 7 citations
- Temporal-Aware Evaluation and Learning for Temporal Graph Neural NetworksJunwei Su, Shan WuAAAI 2025 · 3 citations
- Unlocking Multi-Modal Potentials for Link Prediction on Dynamic Text-Attributed GraphsYuanyuan Xu, Wenjie Zhang, Ying Zhang, Xuemin Lin et al.AAAI 2026 · 2 citations
- TIDFormer: Exploiting Temporal and Interactive Dynamics Makes A Great Dynamic Graph TransformerJie Peng, Zhewei Wei, Yuhang YeKDD 2025 · 2 citations
Builds on14
- EvolveGCN: Evolving Graph Convolutional Networks for Dynamic GraphsAldo Pareja, Giacomo Domeniconi, Jie Chen, Tengfei Ma et al.AAAI 2020 · 1,429 citations
- Large Batch Optimization for Deep Learning: Training BERT in 76 minutesYang You, Jing Li, Sashank J. Reddi, Jonathan Hseu et al.ICLR 2020 · 1,170 citations
- Inductive representation learning on temporal graphsDa Xu, Chuanwei Ruan, Evren Körpeoglu, Sushant Kumar et al.ICLR 2020 · 901 citations
- Don't Use Large Mini-batches, Use Local SGDTao Lin, Sebastian U. Stich, Kumar Kshitij Patel, Martin JaggiICLR 2020 · 462 citations
- Is Local SGD Better than Minibatch SGD?Blake E. Woodworth, Kumar Kshitij Patel, Sebastian U. Stich, Zhen Dai et al.ICML 2020 · 277 citations
Related papers
- DistTGL: Distributed Memory-Based Temporal Graph Neural Network TrainingHongkuan Zhou, Da Zheng, Xiang Song, George Karypis et al.SC 2023 · 21 citations
- Effective and Efficient Distributed Temporal Graph Learning through Hotspot Memory SharingLongjiao Zhang, Rui Wang, Tongya Zheng, Ziqi Huang et al.VLDB 2025 · 1 citation
- SEIGN: A Simple and Efficient Graph Neural Network for Large Dynamic GraphsXiao Qin, Nasrullah Sheikh, Chuan Lei, Berthold Reinwald et al.ICDE 2023 · 14 citations
- PipeTGL: (Near) Zero Bubble Memory-based Temporal Graph Neural Network Training via Pipeline OptimizationJun Liu, Bingqian Du, Ziyue Luo, Sitian Lu et al.VLDB 2025
- PiPAD: Pipelined and Parallel Dynamic GNN Training on GPUsChunyang Wang, Desen Sun, Yuebin BaiPPoPP 2023 · 27 citations
