Effective and Efficient Distributed Temporal Graph Learning through Hotspot Memory Sharing
Longjiao Zhang, Rui Wang, Tongya Zheng, Ziqi Huang, Wenjie Huang, Xinyu Wang, Can Wang, Mingli Song, Sai Wu, Shuibing He
Abstract
Memory-based temporal graph neural network (MTGNN) models are effective for predicting temporal graphs by using node memory and message-passing modules to capture temporal and structural information, respectively. However, distributed training for large graphs presents challenges such as accuracy loss and decreased efficiency due to remote features and memory transmission. Despite improvements in MTGNN system optimizations, issues like dynamic load imbalances, communication overhead, and memory staleness persist. To tackle these challenges, we introduce MemShare, a distributed MTGNN system. MemShare introduces a novel shared node memory paradigm that utilizes a small subset of shared nodes across machines and GPUs to reduce distributed communication for memory management. It incorporates techniques like shared nodes-centric graph partitioning, shared nodes-aware boundary decay sampling, and shared nodes-targeted synchronous smoothing aggregation. Experiments show that MemShare outperforms existing distributed MTGNN systems in accuracy and training efficiency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4e9a6153-2105-48e3-900f-936ba075eeaeCited by top-tier papers2
- Neural Graph Navigation for Intelligent Subgraph MatchingYuchen Ying, Yiyang Dai, Wenda Li, Wenjie Huang et al.AAAI 2026
- FlareDTDG: Harnessing Temporal Recency for Scalable Discrete-Time Dynamic Graph TrainingWenjie Huang, Rui Wang, Jing Cao, Tongya Zheng et al.VLDB 2026
Builds on22
- Inductive representation learning on temporal graphsDa Xu, Chuanwei Ruan, Evren Körpeoglu, Sushant Kumar et al.ICLR 2020 · 901 citations
- Inductive Representation Learning in Temporal Networks via Causal Anonymous WalksYanbang Wang, Yen-Yu Chang, Yunyu Liu, Jure Leskovec et al.ICLR 2021 · 326 citations
- DistGNN: scalable distributed training for large-scale graph neural networksMd. Vasimuddin, Sanchit Misra, Guixiang Ma, Ramanarayan Mohanty et al.SC 2021 · 110 citations
- TGL: A General Framework for Temporal GNN Training onBillion-Scale GraphsHongkuan Zhou, Da Zheng, Israt Nisa, Vassilis N. Ioannidis et al.VLDB 2022 · 109 citations
- ByteGNN: Efficient Graph Neural Network Training at Large ScaleChenguang Zheng, Hongzhi Chen, Yuxuan Cheng, Zhezheng Song et al.VLDB 2022 · 107 citations
Related papers
- DistTGL: Distributed Memory-Based Temporal Graph Neural Network TrainingHongkuan Zhou, Da Zheng, Xiang Song, George Karypis et al.SC 2023 · 21 citations
- MSPipe: Efficient Temporal GNN Training via Staleness-Aware PipelineGuangming Sheng, Junwei Su, Chao Huang, Chuan WuKDD 2024 · 7 citations
- Efficient scaling of dynamic graph neural networksVenkatesan T. Chakaravarthy, Shivmaran S. Pandian, Saurabh Raje, Yogish Sabharwal et al.SC 2021 · 35 citations
- TimeSGN: Scalable and Effective Temporal Graph Neural NetworkYuanyuan Xu, Wenjie Zhang, Ying Zhang, Maria E. Orlowska et al.ICDE 2024 · 15 citations
- PRES: Toward Scalable Memory-Based Dynamic Graph Neural NetworksJunwei Su, Difan Zou, Chuan WuICLR 2024 · 14 citations
