Convolutional Tensor-Train LSTM for Spatio-Temporal Learning
Jiahao Su, Wonmin Byeon, Jean Kossaifi, Furong Huang, Jan Kautz, Anima Anandkumar
摘要
Learning from spatio-temporal data has numerous applications such as humanbehavior analysis, object tracking, video compression, and physics simulation. However, existing methods still perform poorly on challenging video tasks such as long-term forecasting. This is because these kinds of challenging tasks require learning long-term spatio-temporal correlations in the video sequence. In this paper, we propose a higher-order convolutional LSTM model that can efficiently learn these correlations, along with a succinct representations of the history. This is accomplished through a novel tensor-train module that performs prediction by combining convolutional features across time. To make this feasible in terms of computation and memory requirements, we propose a novel convolutional tensortrain decomposition of the higher-order model. This decomposition reduces the model complexity by jointly approximating a sequence of convolutional kernels as a low-rank tensor-train factorization. As a result, our model outperforms existing approaches, but uses only a fraction of parameters, including the baseline models. Our results achieve state-of-the-art performance in a wide range of applications and datasets, including the multi-steps video prediction on the Moving-MNIST-2 and KTH action datasets as well as early activity recognition on the Something-Something V2 dataset. Learning from (video) sequences. Most state-of-the-art video models are based on recurrent neural networks (RNNs), typically some variations of Convolutional LSTM (ConvLSTM) where spatiotemporal information is encoded explicitly in each cell [4] [5] [6] [7] . These RNNs are first-order Markovian * Equal contribution † This work was done while the author was an intern at NVIDIA. Project page: https://sites.google.com/nvidia.com/conv-tt-lstm Preprint. Under review.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Convolutional State Space Models for Long-Range Spatiotemporal ModelingJimmy T. H. Smith, Shalini De Mello, Jan Kautz, Scott W. Linderman 等NeurIPS 2023 · 被引用 36 次
- Extrapolation and Spectral Bias of Neural Nets with Hadamard Product: a Polynomial Net StudyYongtao Wu, Zhenyu Zhu, Fanghui Liu, Grigorios Chrysos 等NeurIPS 2022 · 被引用 19 次
- MIMO Is All You Need:A Strong Multi-in-Multi-Out Baseline for Video PredictionShuliang Ning, Mengcheng Lan, Yanran Li, Chaofeng Chen 等AAAI 2023 · 被引用 14 次
- Controlling the Complexity and Lipschitz Constant improves Polynomial NetsZhenyu Zhu, Fabian Latorre, Grigorios Chrysos, Volkan CevherICLR 2022 · 被引用 12 次
- Ditto: Accelerating Diffusion Model via Temporal Value SimilaritySungbin Kim, Hyunwuk Lee, Wonho Cho, Mincheol Park 等HPCA 2025 · 被引用 9 次
它引用的顶会 Paper2
相关 Paper
- Towards Extremely Compact RNNs for Video Recognition With Fully Decomposed Hierarchical Tucker StructureMiao Yin, Siyu Liao, Xiao-Yang Liu, Xiaodong Wang 等CVPR 2021
- CT-Net: Channel Tensorization Network for Video ClassificationKunchang Li, Xianhang Li, Yali Wang, Jun Wang 等ICLR 2021 · 被引用 69 次
- Deep Learning in Latent Space for Video Prediction and CompressionBowen Liu, Yu Chen, Shiyu Liu, Hun-Seok KimCVPR 2021
- Towards Efficient Tensor Decomposition-Based DNN Model Compression With Optimization FrameworkMiao Yin, Yang Sui, Siyu Liao, Bo YuanCVPR 2021
- SSAN: Separable Self-Attention Network for Video Representation LearningXudong Guo, Xun Guo, Yan LuCVPR 2021
