SwinLSTM: Improving Spatiotemporal Prediction Accuracy using Swin Transformer and LSTM
Song Tang, Chuang Li, Pu Zhang, Rongnian Tang
摘要
Integrating CNNs and RNNs to capture spatiotemporal dependencies is a prevalent strategy for spatiotemporal prediction tasks. However, the property of CNNs to learn local spatial information decreases their efficiency in capturing spatiotemporal dependencies, thereby limiting their prediction accuracy. In this paper, we propose a new recurrent cell, SwinLSTM, which integrates Swin Transformer blocks and the simplified LSTM, an extension that replaces the convolutional structure in ConvLSTM with the self-attention mechanism. Furthermore, we construct a network with SwinLSTM cell as the core for spatiotemporal prediction. Without using unique tricks, SwinLSTM outperforms state-of-the-art methods on Moving MNIST, Human3.6m, TaxiBJ, and KTH datasets. In particular, it exhibits a significant improvement in prediction accuracy compared to ConvLSTM. Our competitive experimental results demonstrate that learning global spatial dependencies is more advantageous for models to capture spatiotemporal dependencies. We hope that SwinLSTM can serve as a solid baseline to promote the advancement of spatiotemporal prediction accuracy. The codes are publicly available at https://github.com/SongTang-x/SwinLSTM .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Advection Augmented Convolutional Neural NetworksNiloufar Zakariaei, Siddharth Rout, Eldad Haber, Moshe EliasofNeurIPS 2024 · 被引用 7 次
- PFGNet: A Fully Convolutional Frequency-Guided Peripheral Gating Network for Efficient Spatiotemporal Predictive LearningXinyong Cai, Changbin Sun, Yong Wang, Hongyu Yang 等CVPR 2026 · 被引用 3 次
- DFDNet: Disentangling and Filtering Dynamics for Enhanced Video PredictionLianqiang Gan, Junyu Lai, Jingze Ju, Lianli Gao 等AAAI 2025 · 被引用 2 次
- -Net: A Physics-Informed Spatio-Temporal Model for Global Surface ReconstructionHao Zheng, Yuting Zheng, Hanbo Huang, Chaofan Sun 等ICCV 2025
- Online Generic Event Boundary DetectionHyungrok Jung, Daneul Kim, Seunggyun Lim, Jeany Son 等ICCV 2025
它引用的顶会 Paper14
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar 等NeurIPS 2021 · 被引用 9,661 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- Multi-Scale Vision Longformer: A New Vision Transformer for High-Resolution Image EncodingPengchuan Zhang, Xiyang Dai, Jianwei Yang, Bin Xiao 等ICCV 2021 · 被引用 384 次
相关 Paper
- Self-Attention ConvLSTM for Spatiotemporal PredictionZhihui Lin, Maomao Li, Zhuobin Zheng, Yangyang Cheng 等AAAI 2020 · 被引用 347 次
- Modeling Citywide Crowd Flows using Attentive Convolutional LSTMChi Harold Liu, Chengzhe Piao, Xiaoxin Ma, Ye Yuan 等ICDE 2021 · 被引用 24 次
- Swin-UNIT: Transformer-based GAN for High-resolution Unpaired Image TranslationYifan Li, Yaochen Li, Wenneng Tang, Zhifeng Zhu 等ACM MM 2023 · 被引用 13 次
- Learning Trajectory Dependencies for Human Motion PredictionWei Mao, Miaomiao Liu, Mathieu Salzmann, Hongdong LiICCV 2019 · 被引用 534 次
- SwinTrack: A Simple and Strong Baseline for Transformer TrackingLiting Lin, Heng Fan, Zhipeng Zhang, Yong Xu 等NeurIPS 2022 · 被引用 556 次
