Temporal Attention Unit: Towards Efficient Spatiotemporal Predictive Learning
Cheng Tan, Zhangyang Gao, Lirong Wu, Yongjie Xu, Jun Xia, Siyuan Li, Stan Z. Li
摘要
Spatiotemporal predictive learning aims to generate future frames by learning from historical frames. In this paper, we investigate existing methods and present a general framework of spatiotemporal predictive learning, in which the spatial encoder and decoder capture intra-frame features and the middle temporal module catches inter-frame correlations. While the mainstream methods employ recurrent units to capture long-term temporal dependencies, they suffer from low computational efficiency due to their unparallelizable architectures. To parallelize the temporal module, we propose the Temporal Attention Unit (TAU), which decomposes temporal attention into intra-frame statical attention and inter-frame dynamical attention. Moreover, while the mean squared error loss focuses on intra-frame errors, we introduce a novel differential divergence regularization to take inter-frame variations into account. Extensive experiments demonstrate that the proposed method enables the derived model to achieve competitive performance on various spatiotemporal prediction benchmarks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper30
- UniST: A Prompt-Empowered Universal Model for Urban Spatio-Temporal PredictionYuan Yuan, Jingtao Ding, Jie Feng, Depeng Jin 等KDD 2024 · 被引用 75 次
- DiffCast: A Unified Framework via Residual Diffusion for Precipitation NowcastingDemin Yu, Xutao Li, Yunming Ye, Baoquan Zhang 等CVPR 2024 · 被引用 45 次
- NuwaDynamics: Discovering and Updating in Causal Spatio-Temporal ModelingKun Wang, Hao Wu, Yifan Duan, Guibin Zhang 等ICLR 2024 · 被引用 38 次
- Fourier Amplitude and Correlation Loss: Beyond Using L2 Loss for Skillful Precipitation NowcastingChiu Wai Yan, Shi Quan Foo, Van-Hoan Trinh, Dit-Yan Yeung 等NeurIPS 2024 · 被引用 29 次
- Diffusion Transformers as Open-World Spatiotemporal Foundation ModelsYuan Yuan, Chonghua Han, Jingtao Ding, Guozhen Zhang 等NeurIPS 2025 · 被引用 16 次
它引用的顶会 Paper28
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer 等CVPR 2022 · 被引用 6,782 次
相关 Paper
- MAU: A Motion-Aware Unit for Video Prediction and BeyondZheng Chang, Xinfeng Zhang, Shanshe Wang, Siwei Ma 等NeurIPS 2021 · 被引用 193 次
- Wavelet-Driven Spatiotemporal Predictive Learning: Bridging Frequency and Time VariationsXuesong Nie, Yunfeng Yan, Siyuan Li, Cheng Tan 等AAAI 2024 · 被引用 31 次
- ExtDM: Distribution Extrapolation Diffusion Model for Video PredictionZhicheng Zhang, Junyao Hu, Wentao Cheng, Danda Pani Paudel 等CVPR 2024 · 被引用 24 次
- Dynamical Diffusion: Learning Temporal Dynamics with Diffusion ModelsXingzhuo Guo, Yu Zhang, Baixu Chen, Haoran Xu 等ICLR 2025
- Predictive Feature Learning for Future Segmentation PredictionZihang Lin, Jiangxin Sun, Jianfang Hu, Qi-Zhi Yu 等ICCV 2021 · 被引用 18 次
