SwinLSTM: Improving Spatiotemporal Prediction Accuracy using Swin Transformer and LSTM
Song Tang, Chuang Li, Pu Zhang, Rongnian Tang
Abstract
Integrating CNNs and RNNs to capture spatiotemporal dependencies is a prevalent strategy for spatiotemporal prediction tasks. However, the property of CNNs to learn local spatial information decreases their efficiency in capturing spatiotemporal dependencies, thereby limiting their prediction accuracy. In this paper, we propose a new recurrent cell, SwinLSTM, which integrates Swin Transformer blocks and the simplified LSTM, an extension that replaces the convolutional structure in ConvLSTM with the self-attention mechanism. Furthermore, we construct a network with SwinLSTM cell as the core for spatiotemporal prediction. Without using unique tricks, SwinLSTM outperforms state-of-the-art methods on Moving MNIST, Human3.6m, TaxiBJ, and KTH datasets. In particular, it exhibits a significant improvement in prediction accuracy compared to ConvLSTM. Our competitive experimental results demonstrate that learning global spatial dependencies is more advantageous for models to capture spatiotemporal dependencies. We hope that SwinLSTM can serve as a solid baseline to promote the advancement of spatiotemporal prediction accuracy. The codes are publicly available at https://github.com/SongTang-x/SwinLSTM .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers11
- Advection Augmented Convolutional Neural NetworksNiloufar Zakariaei, Siddharth Rout, Eldad Haber, Moshe EliasofNeurIPS 2024 · 7 citations
- PFGNet: A Fully Convolutional Frequency-Guided Peripheral Gating Network for Efficient Spatiotemporal Predictive LearningXinyong Cai, Changbin Sun, Yong Wang, Hongyu Yang et al.CVPR 2026 · 3 citations
- DFDNet: Disentangling and Filtering Dynamics for Enhanced Video PredictionLianqiang Gan, Junyu Lai, Jingze Ju, Lianli Gao et al.AAAI 2025 · 2 citations
- -Net: A Physics-Informed Spatio-Temporal Model for Global Surface ReconstructionHao Zheng, Yuting Zheng, Hanbo Huang, Chaofan Sun et al.ICCV 2025
- Online Generic Event Boundary DetectionHyungrok Jung, Daneul Kim, Seunggyun Lim, Jeany Son et al.ICCV 2025
Builds on14
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Multi-Scale Vision Longformer: A New Vision Transformer for High-Resolution Image EncodingPengchuan Zhang, Xiyang Dai, Jianwei Yang, Bin Xiao et al.ICCV 2021 · 384 citations
Related papers
- Self-Attention ConvLSTM for Spatiotemporal PredictionZhihui Lin, Maomao Li, Zhuobin Zheng, Yangyang Cheng et al.AAAI 2020 · 347 citations
- Modeling Citywide Crowd Flows using Attentive Convolutional LSTMChi Harold Liu, Chengzhe Piao, Xiaoxin Ma, Ye Yuan et al.ICDE 2021 · 24 citations
- Swin-UNIT: Transformer-based GAN for High-resolution Unpaired Image TranslationYifan Li, Yaochen Li, Wenneng Tang, Zhifeng Zhu et al.ACM MM 2023 · 13 citations
- Learning Trajectory Dependencies for Human Motion PredictionWei Mao, Miaomiao Liu, Mathieu Salzmann, Hongdong LiICCV 2019 · 534 citations
- SwinTrack: A Simple and Strong Baseline for Transformer TrackingLiting Lin, Heng Fan, Zhipeng Zhang, Yong Xu et al.NeurIPS 2022 · 556 citations
