Convolutional Tensor-Train LSTM for Spatio-Temporal Learning
Jiahao Su, Wonmin Byeon, Jean Kossaifi, Furong Huang, Jan Kautz, Anima Anandkumar
Abstract
Learning from spatio-temporal data has numerous applications such as humanbehavior analysis, object tracking, video compression, and physics simulation. However, existing methods still perform poorly on challenging video tasks such as long-term forecasting. This is because these kinds of challenging tasks require learning long-term spatio-temporal correlations in the video sequence. In this paper, we propose a higher-order convolutional LSTM model that can efficiently learn these correlations, along with a succinct representations of the history. This is accomplished through a novel tensor-train module that performs prediction by combining convolutional features across time. To make this feasible in terms of computation and memory requirements, we propose a novel convolutional tensortrain decomposition of the higher-order model. This decomposition reduces the model complexity by jointly approximating a sequence of convolutional kernels as a low-rank tensor-train factorization. As a result, our model outperforms existing approaches, but uses only a fraction of parameters, including the baseline models. Our results achieve state-of-the-art performance in a wide range of applications and datasets, including the multi-steps video prediction on the Moving-MNIST-2 and KTH action datasets as well as early activity recognition on the Something-Something V2 dataset. Learning from (video) sequences. Most state-of-the-art video models are based on recurrent neural networks (RNNs), typically some variations of Convolutional LSTM (ConvLSTM) where spatiotemporal information is encoded explicitly in each cell [4] [5] [6] [7] . These RNNs are first-order Markovian * Equal contribution † This work was done while the author was an intern at NVIDIA. Project page: https://sites.google.com/nvidia.com/conv-tt-lstm Preprint. Under review.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b52743fe-92f9-4ed8-a67a-240b1ea1ac7eCited by top-tier papers14
- Convolutional State Space Models for Long-Range Spatiotemporal ModelingJimmy T. H. Smith, Shalini De Mello, Jan Kautz, Scott W. Linderman et al.NeurIPS 2023 · 36 citations
- Extrapolation and Spectral Bias of Neural Nets with Hadamard Product: a Polynomial Net StudyYongtao Wu, Zhenyu Zhu, Fanghui Liu, Grigorios Chrysos et al.NeurIPS 2022 · 19 citations
- MIMO Is All You Need:A Strong Multi-in-Multi-Out Baseline for Video PredictionShuliang Ning, Mengcheng Lan, Yanran Li, Chaofeng Chen et al.AAAI 2023 · 14 citations
- Controlling the Complexity and Lipschitz Constant improves Polynomial NetsZhenyu Zhu, Fabian Latorre, Grigorios Chrysos, Volkan CevherICLR 2022 · 12 citations
- Ditto: Accelerating Diffusion Model via Temporal Value SimilaritySungbin Kim, Hyunwuk Lee, Wonho Cho, Mincheol Park et al.HPCA 2025 · 9 citations
Builds on2
Related papers
- Towards Extremely Compact RNNs for Video Recognition With Fully Decomposed Hierarchical Tucker StructureMiao Yin, Siyu Liao, Xiao-Yang Liu, Xiaodong Wang et al.CVPR 2021
- CT-Net: Channel Tensorization Network for Video ClassificationKunchang Li, Xianhang Li, Yali Wang, Jun Wang et al.ICLR 2021 · 69 citations
- Deep Learning in Latent Space for Video Prediction and CompressionBowen Liu, Yu Chen, Shiyu Liu, Hun-Seok KimCVPR 2021
- Towards Efficient Tensor Decomposition-Based DNN Model Compression With Optimization FrameworkMiao Yin, Yang Sui, Siyu Liao, Bo YuanCVPR 2021
- SSAN: Separable Self-Attention Network for Video Representation LearningXudong Guo, Xun Guo, Yan LuCVPR 2021
