Self-Attention ConvLSTM for Spatiotemporal Prediction
Zhihui Lin, Maomao Li, Zhuobin Zheng, Yangyang Cheng, Chun Yuan
Abstract
Spatiotemporal prediction is challenging due to the complex dynamic motion and appearance changes. Existing work concentrates on embedding additional cells into the standard ConvLSTM to memorize spatial appearances during the prediction. These models always rely on the convolution layers to capture the spatial dependence, which are local and inefficient. However, long-range spatial dependencies are significant for spatial applications. To extract spatial features with both global and local dependencies, we introduce the self-attention mechanism into ConvLSTM. Specifically, a novel self-attention memory (SAM) is proposed to memorize features with long-range dependencies in terms of spatial and temporal domains. Based on the self-attention, SAM can produce features by aggregating features across all positions of both the input itself and memory features with pair-wise similarity scores. Moreover, the additional memory is updated by a gating mechanism on aggregated features and an established highway with the memory of the previous time step. Therefore, through SAM, we can extract features with long-range spatiotemporal dependencies. Furthermore, we embed the SAM into a standard ConvLSTM to construct a self-attention ConvLSTM (SA-ConvLSTM) for the spatiotemporal prediction. In experiments, we apply the SA-ConvLSTM to perform frame prediction on the MovingMNIST and KTH datasets and traffic flow prediction on the TexiBJ dataset. Our SA-ConvLSTM achieves state-of-the-art results on both datasets with fewer parameters and higher time efficiency than previous state-of-the-art method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- PDFormer: Propagation Delay-Aware Dynamic Long-Range Transformer for Traffic Flow PredictionJiawei Jiang, Chengkai Han, Wayne Xin Zhao, Jingyuan WangAAAI 2023 · 542 citations
- UniST: A Prompt-Empowered Universal Model for Urban Spatio-Temporal PredictionYuan Yuan, Jingtao Ding, Jie Feng, Depeng Jin et al.KDD 2024 · 75 citations
- Iso-Dream: Isolating and Leveraging Noncontrollable Visual Dynamics in World ModelsMinting Pan, Xiangming Zhu, Yunbo Wang, Xiaokang YangNeurIPS 2022 · 74 citations
- PastNet: Introducing Physical Inductive Biases for Spatio-temporal Video PredictionHao Wu, Fan Xu, Chong Chen, Xian-Sheng Hua et al.ACM MM 2024 · 36 citations
- Diffusion Transformers as Open-World Spatiotemporal Foundation ModelsYuan Yuan, Chonghua Han, Jingtao Ding, Guozhen Zhang et al.NeurIPS 2025 · 16 citations
Builds on1
Related papers
- SwinLSTM: Improving Spatiotemporal Prediction Accuracy using Swin Transformer and LSTMSong Tang, Chuang Li, Pu Zhang, Rongnian TangICCV 2023 · 117 citations
- Modeling Citywide Crowd Flows using Attentive Convolutional LSTMChi Harold Liu, Chengzhe Piao, Xiaoxin Ma, Ye Yuan et al.ICDE 2021 · 24 citations
- ST-ABC: Spatio-Temporal Attention-Based Convolutional Network for Multi-Scale Lane-Level Traffic PredictionShuhao Li, Yue Cui, Libin Li, Weidong Yang et al.ICDE 2024 · 13 citations
- Convolutional State Space Models for Long-Range Spatiotemporal ModelingJimmy T. H. Smith, Shalini De Mello, Jan Kautz, Scott W. Linderman et al.NeurIPS 2023 · 36 citations
- SSL-STMFormer Self-Supervised Learning Spatio-Temporal Entanglement Transformer for Traffic Flow PredictionZetao Li, Zheng Hu, Peng Han, Yu Gu et al.AAAI 2025 · 12 citations
