Convolutional State Space Models for Long-Range Spatiotemporal Modeling
Jimmy T. H. Smith, Shalini De Mello, Jan Kautz, Scott W. Linderman, Wonmin Byeon
Abstract
Effectively modeling long spatiotemporal sequences is challenging due to the need to model complex spatial correlations and long-range temporal dependencies simultaneously. ConvLSTMs attempt to address this by updating tensor-valued states with recurrent neural networks, but their sequential computation makes them slow to train. In contrast, Transformers can process an entire spatiotemporal sequence, compressed into tokens, in parallel. However, the cost of attention scales quadratically in length, limiting their scalability to longer sequences. Here, we address the challenges of prior methods and introduce convolutional state space models (ConvSSM) 1 that combine the tensor modeling ideas of ConvLSTM with the long sequence modeling approaches of state space methods such as S4 and S5. First, we demonstrate how parallel scans can be applied to convolutional recurrences to achieve subquadratic parallelization and fast autoregressive generation. We then establish an equivalence between the dynamics of ConvSSMs and SSMs, which motivates parameterization and initialization strategies for modeling long-range dependencies. The result is ConvS5, an efficient ConvSSM variant for long-range spatiotemporal modeling. ConvS5 significantly outperforms Transformers and ConvLSTM on a long horizon Moving-MNIST experiment while training 3× faster than ConvLSTM and generating samples 400× faster than Transformers. In addition, ConvS5 matches or exceeds the performance of state-of-the-art methods on challenging DMLab, Minecraft and Habitat prediction benchmarks and enables new directions for modeling long spatiotemporal sequences.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space ModelLianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang et al.ICML 2024 · 1,725 citations
- Quanta Neural Networks: From Photons to PerceptionVarun Sundar, Tianyi Zhang, Sacha Jungerman, Mohit GuptaICCV 2025 · 1 citation
- How2Compress: Scalable and Efficient Edge Video Analytics via Adaptive Granular Video CompressionYuheng Wu, Thanh-Tung Nguyen, Lucas Liebe, Quang Tau et al.ACM MM 2025 · 1 citation
- Explicit Context Reasoning with Supervision for Visual TrackingFansheng Zeng, Bineng Zhong, Haiying Xia, Yufei Tan et al.ACM MM 2025 · 1 citation
- Tuning Frequency Bias of State Space ModelsAnnan Yu, Dongwei Lyu, Soon Hoe Lim, Michael W. Mahoney et al.ICLR 2025
Builds on43
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan et al.NeurIPS 2022 · 2,948 citations
Related papers
- ConvT3: Structured State Kernels for Convolutional State Space ModelsJaeyoung Hong, Yun Young Choi, Joohwan Ko, Minseon GwakICLR 2026
- VGA: Hardware Accelerator for Scalable Long Sequence Model InferenceSeung Yul Lee, Hyunseung Lee, Jihoon Hong, SangLyul Cho et al.MICRO 2024 · 7 citations
- Simplified State Space Layers for Sequence ModelingJimmy T. H. Smith, Andrew Warrington, Scott W. LindermanICLR 2023 · 78 citations
- Self-Attention ConvLSTM for Spatiotemporal PredictionZhihui Lin, Maomao Li, Zhuobin Zheng, Yangyang Cheng et al.AAAI 2020 · 347 citations
- Block-State TransformersJonathan Pilault, Mahan Fathi, Orhan Firat, Chris Pal et al.NeurIPS 2023 · 33 citations
