MAU: A Motion-Aware Unit for Video Prediction and Beyond
Zheng Chang, Xinfeng Zhang, Shanshe Wang, Siwei Ma, Yan Ye, Xiang Xinguang, Wen Gao
摘要
Accurately predicting inter-frame motion information plays a key role in video prediction tasks. In this paper, we propose a Motion-Aware Unit (MAU) to capture reliable inter-frame motion information by broadening the temporal receptive field of the predictive units. The MAU consists of two modules, the attention module and the fusion module. The attention module aims to learn an attention map based on the correlations between the current spatial state and the historical spatial states. Based on the learned attention map, the historical temporal states are aggregated to an augmented motion information (AMI). In this way, the predictive unit can perceive more temporal dynamics from a wider receptive field. Then, the fusion module is utilized to further aggregate the augmented motion information (AMI) and current appearance information (current spatial state) to the final predicted frame. The computation load of MAU is relatively low, and the proposed unit can be easily applied to other predictive models. Moreover, an information recalling scheme is employed into the encoders and decoders to help preserve the visual details of the predictions. We evaluate the MAU on both video prediction and early action recognition tasks. Experimental results show that the MAU outperforms the state-of-the-art methods on both tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- SwinLSTM: Improving Spatiotemporal Prediction Accuracy using Swin Transformer and LSTMSong Tang, Chuang Li, Pu Zhang, Rongnian TangICCV 2023 · 被引用 117 次
- UniST: A Prompt-Empowered Universal Model for Urban Spatio-Temporal PredictionYuan Yuan, Jingtao Ding, Jie Feng, Depeng Jin 等KDD 2024 · 被引用 75 次
- Earthfarsser: Versatile Spatio-Temporal Dynamical Systems Modeling in One ModelHao Wu, Yuxuan Liang, Wei Xiong, Zhengyang Zhou 等AAAI 2024 · 被引用 58 次
- DiffCast: A Unified Framework via Residual Diffusion for Precipitation NowcastingDemin Yu, Xutao Li, Yunming Ye, Baoquan Zhang 等CVPR 2024 · 被引用 45 次
- Wavelet-Driven Spatiotemporal Predictive Learning: Bridging Frequency and Time VariationsXuesong Nie, Yunfeng Yan, Siyuan Li, Cheng Tan 等AAAI 2024 · 被引用 31 次
它引用的顶会 Paper5
- Stochastic Latent Residual Video PredictionJean-Yves Franceschi, Edouard Delasalles, Mickaël Chen, Sylvain Lamprier 等ICML 2020 · 被引用 166 次
- Efficient and Information-Preserving Future Frame Prediction and BeyondWei Yu, Yichao Lu, Steve Easterbrook, Sanja FidlerICLR 2020 · 被引用 127 次
- Exploring Spatial-Temporal Multi-Frequency Analysis for High-Fidelity and Temporal-Consistency Video PredictionBeibei Jin, Yu Hu, Qiankun Tang, Jingyu Niu 等CVPR 2020
- Greedy Hierarchical Variational Autoencoders for Large-Scale Video PredictionBohan Wu, Suraj Nair, Roberto Martín-Martín, Li Fei-Fei 等CVPR 2021
- MotionRNN: A Flexible Model for Video Prediction With Spacetime-Varying MotionsHaixu Wu, Zhiyu Yao, Jianmin Wang, Mingsheng LongCVPR 2021
相关 Paper
- Temporal Attention Unit: Towards Efficient Spatiotemporal Predictive LearningCheng Tan, Zhangyang Gao, Lirong Wu, Yongjie Xu 等CVPR 2023
- ExtDM: Distribution Extrapolation Diffusion Model for Video PredictionZhicheng Zhang, Junyao Hu, Wentao Cheng, Danda Pani Paudel 等CVPR 2024 · 被引用 24 次
- TEA: Temporal Excitation and Aggregation for Action RecognitionYan Li, Bin Ji, Xintian Shi, Jianguo Zhang 等CVPR 2020
- Extracting Motion and Appearance via Inter-Frame Attention for Efficient Video Frame InterpolationGuozhen Zhang, Yuhan Zhu, Haonan Wang, Youxin Chen 等CVPR 2023
- Motion Guided Region Message Passing for Video CaptioningShaoxiang Chen, Yu-Gang JiangICCV 2021 · 被引用 71 次
