SC2025Top-tier venue
PGT-I: Scaling Spatiotemporal GNNs with Memory-Efficient Distributed Training
Seth Ockerman, Amal Gueroudji, Tanwi Mallick, Yixuan He, Line Pouchard, Robert B. Ross, Shivaram Venkataraman
Abstract
Spatiotemporal graph neural networks (ST-GNNs) are powerful tools for modeling spatial and temporal data dependencies. However, their applications have been limited primarily to small-scale datasets because of memory constraints. While distributed training offers a solution, current frameworks lack support for spatiotemporal models and overlook the properties of spatiotemporal data. Informed by a scaling study on a large-scale workload, we present PyTorch Geometric Temporal Index (PGT-I), an extension to Py-Torch Geometric Temporal that integrates distributed data parallel training and two novel strategies: index-batching and distributedindex-batching. Our index techniques exploit spatiotemporal structure to construct snapshots dynamically at runtime, significantly reducing memory overhead, while distributed-index-batching extends this approach by enabling scalable processing across multiple GPUs. Our techniques enable the first-ever training of an ST-GNN on the entire PeMS dataset without graph partitioning, reducing peak memory usage by up to 89% and achieving up to a 11.78x speedup over standard DDP with 128 GPUs. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on6
- Adaptive Graph Convolutional Recurrent Network for Traffic ForecastingLei Bai, Lina Yao, Can Li, Xianzhi Wang et al.NeurIPS 2020 · 2,206 citations
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 852 citations
- Conditional Local Convolution for Spatio-Temporal Meteorological ForecastingHaitao Lin, Zhangyang Gao, Yongjie Xu, Lirong Wu et al.AAAI 2022 · 116 citations
- Learning the Evolutionary and Multi-scale Graph Structure for Multivariate Time Series ForecastingJunchen Ye, Zihan Liu, Bowen Du, Leilei Sun et al.KDD 2022 · 109 citations
- DistTGL: Distributed Memory-Based Temporal Graph Neural Network TrainingHongkuan Zhou, Da Zheng, Xiang Song, George Karypis et al.SC 2023 · 21 citations
Related papers
- SWIFT: Enabling Large-Scale Temporal Graph Learning on a Single MachineRui Guo, Zezhong Ding, Xike Xie, Jianliang XuSIGMOD 2026 · 2 citations
- WholeGraph: A Fast Graph Neural Network Training Framework with Multi-GPU Distributed Shared Memory ArchitectureDongxu Yang, Junhong Liu, Jiaxing Qi, Junjie LaiSC 2022 · 12 citations
- DGC: Training Dynamic Graphs with Spatio-Temporal Non-Uniformity using Graph Partitioning by ChunksFahao Chen, Peng Li, Celimuge WuSIGMOD 2024 · 10 citations
- GNNAutoScale: Scalable and Expressive Graph Neural Networks via Historical EmbeddingsMatthias Fey, Jan Eric Lenssen, Frank Weichert, Jure LeskovecICML 2021 · 149 citations
- P3: Distributed Deep Graph Learning at ScaleSwapnil Gandhi, Anand Padmanabha IyerOSDI 2021 · 192 citations
