MSPipe: Efficient Temporal GNN Training via Staleness-Aware Pipeline
Guangming Sheng, Junwei Su, Chao Huang, Chuan Wu
摘要
Memory-based Temporal Graph Neural Networks (MTGNNs) are a class of temporal graph neural networks that utilize a node memory module to capture and retain long-term temporal dependencies, leading to superior performance compared to memory-less counterparts. However, the iterative reading and updating process of the memory module in MTGNNs to obtain up-to-date information needs to follow the temporal dependencies. This introduces significant overhead and limits training throughput. Existing optimizations for static GNNs are not directly applicable to MTGNNs due to differences in training paradigm, model architecture, and the absence of a memory module. Moreover, these optimizations do not effectively address the challenges posed by temporal dependencies, making them ineffective for MTGNN training. In this paper, we propose MSPipe, a general and efficient framework for memory-based TGNNs that maximizes training throughput while maintaining model accuracy. Our design specifically addresses the unique challenges associated with fetching and updating node memory states in MTGNNs by integrating staleness into the memory module. However, simply introducing a predefined staleness bound in the memory module to break temporal dependencies may lead to suboptimal performance and lack of generalizability across different models and datasets. To overcome this, we introduce an online pipeline scheduling algorithm in MSPipe that strategically breaks temporal dependencies with minimal staleness and delays memory fetching to obtain fresher memory states. This is achieved without stalling the MTGNN training stage or causing resource contention. Additionally, we design a staleness mitigation mechanism to enhance training convergence and model accuracy. Furthermore, we provide convergence analysis and demonstrate that MSPipe maintains the same convergence rate as vanilla sampling-based GNN training. Experimental results show that MSPipe achieves up to 2.45× speed-up without sacrificing accuracy, making it a promising solution for efficient MTGNN training. The implementation of our paper can be found at the following link: https://github.com/PeterSH6/MSPipe.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- PRES: Toward Scalable Memory-Based Dynamic Graph Neural NetworksJunwei Su, Difan Zou, Chuan WuICLR 2024 · 被引用 14 次
- Temporal-Aware Evaluation and Learning for Temporal Graph Neural NetworksJunwei Su, Shan WuAAAI 2025 · 被引用 3 次
- Laminar: A Scalable Asynchronous RL Post-Training FrameworkGuangming Sheng, Yuxuan Tong, Borui Wan, Wang Zhang 等EuroSys 2026 · 被引用 2 次
- Effective and Efficient Distributed Temporal Graph Learning through Hotspot Memory SharingLongjiao Zhang, Rui Wang, Tongya Zheng, Ziqi Huang 等VLDB 2025 · 被引用 1 次
- PRISM: A Training System to Unlock the Potential of Temporal Graph Learning Through Staleness AvoidanceMd Ashraful Islam, Hojae Son, Suhaas Kiran Doddagaddavalli Gangadharaiah, Marco SerafiniVLDB 2026
它引用的顶会 Paper15
- Inductive representation learning on temporal graphsDa Xu, Chuanwei Ruan, Evren Körpeoglu, Sushant Kumar 等ICLR 2020 · 被引用 901 次
- TGL: A General Framework for Temporal GNN Training onBillion-Scale GraphsHongkuan Zhou, Da Zheng, Israt Nisa, Vassilis N. Ioannidis 等VLDB 2022 · 被引用 109 次
- ByteGNN: Efficient Graph Neural Network Training at Large ScaleChenguang Zheng, Hongzhi Chen, Yuxuan Cheng, Zhezheng Song 等VLDB 2022 · 被引用 107 次
- GNNLab: a factored system for sample-based GNN training over GPUsJianbang Yang, Dahai Tang, Xiaoniu Song, Lei Wang 等EuroSys 2022 · 被引用 105 次
- Time Matters: Sequential Recommendation with Complex Temporal InformationWenwen Ye, Shuaiqiang Wang, Xu Chen, Xuepeng Wang 等SIGIR 2020 · 被引用 89 次
相关 Paper
- PipeTGL: (Near) Zero Bubble Memory-based Temporal Graph Neural Network Training via Pipeline OptimizationJun Liu, Bingqian Du, Ziyue Luo, Sitian Lu 等VLDB 2025
- DistTGL: Distributed Memory-Based Temporal Graph Neural Network TrainingHongkuan Zhou, Da Zheng, Xiang Song, George Karypis 等SC 2023 · 被引用 21 次
- PiPAD: Pipelined and Parallel Dynamic GNN Training on GPUsChunyang Wang, Desen Sun, Yuebin BaiPPoPP 2023 · 被引用 27 次
- SWIFT: Enabling Large-Scale Temporal Graph Learning on a Single MachineRui Guo, Zezhong Ding, Xike Xie, Jianliang XuSIGMOD 2026 · 被引用 2 次
- Cascade: A Dependency-aware Efficient Training Framework for Temporal Graph Neural NetworkYue Dai, Xulong Tang, Youtao ZhangASPLOS 2025 · 被引用 4 次
