MAU: A Motion-Aware Unit for Video Prediction and Beyond
Zheng Chang, Xinfeng Zhang, Shanshe Wang, Siwei Ma, Yan Ye, Xiang Xinguang, Wen Gao
Abstract
Accurately predicting inter-frame motion information plays a key role in video prediction tasks. In this paper, we propose a Motion-Aware Unit (MAU) to capture reliable inter-frame motion information by broadening the temporal receptive field of the predictive units. The MAU consists of two modules, the attention module and the fusion module. The attention module aims to learn an attention map based on the correlations between the current spatial state and the historical spatial states. Based on the learned attention map, the historical temporal states are aggregated to an augmented motion information (AMI). In this way, the predictive unit can perceive more temporal dynamics from a wider receptive field. Then, the fusion module is utilized to further aggregate the augmented motion information (AMI) and current appearance information (current spatial state) to the final predicted frame. The computation load of MAU is relatively low, and the proposed unit can be easily applied to other predictive models. Moreover, an information recalling scheme is employed into the encoders and decoders to help preserve the visual details of the predictions. We evaluate the MAU on both video prediction and early action recognition tasks. Experimental results show that the MAU outperforms the state-of-the-art methods on both tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4720164b-5bd1-4ddb-8811-a6df16a21da3Cited by top-tier papers19
- SwinLSTM: Improving Spatiotemporal Prediction Accuracy using Swin Transformer and LSTMSong Tang, Chuang Li, Pu Zhang, Rongnian TangICCV 2023 · 117 citations
- UniST: A Prompt-Empowered Universal Model for Urban Spatio-Temporal PredictionYuan Yuan, Jingtao Ding, Jie Feng, Depeng Jin et al.KDD 2024 · 75 citations
- Earthfarsser: Versatile Spatio-Temporal Dynamical Systems Modeling in One ModelHao Wu, Yuxuan Liang, Wei Xiong, Zhengyang Zhou et al.AAAI 2024 · 58 citations
- DiffCast: A Unified Framework via Residual Diffusion for Precipitation NowcastingDemin Yu, Xutao Li, Yunming Ye, Baoquan Zhang et al.CVPR 2024 · 45 citations
- Wavelet-Driven Spatiotemporal Predictive Learning: Bridging Frequency and Time VariationsXuesong Nie, Yunfeng Yan, Siyuan Li, Cheng Tan et al.AAAI 2024 · 31 citations
Builds on5
- Stochastic Latent Residual Video PredictionJean-Yves Franceschi, Edouard Delasalles, Mickaël Chen, Sylvain Lamprier et al.ICML 2020 · 166 citations
- Efficient and Information-Preserving Future Frame Prediction and BeyondWei Yu, Yichao Lu, Steve Easterbrook, Sanja FidlerICLR 2020 · 127 citations
- Exploring Spatial-Temporal Multi-Frequency Analysis for High-Fidelity and Temporal-Consistency Video PredictionBeibei Jin, Yu Hu, Qiankun Tang, Jingyu Niu et al.CVPR 2020
- Greedy Hierarchical Variational Autoencoders for Large-Scale Video PredictionBohan Wu, Suraj Nair, Roberto Martín-Martín, Li Fei-Fei et al.CVPR 2021
- MotionRNN: A Flexible Model for Video Prediction With Spacetime-Varying MotionsHaixu Wu, Zhiyu Yao, Jianmin Wang, Mingsheng LongCVPR 2021
Related papers
- Temporal Attention Unit: Towards Efficient Spatiotemporal Predictive LearningCheng Tan, Zhangyang Gao, Lirong Wu, Yongjie Xu et al.CVPR 2023
- ExtDM: Distribution Extrapolation Diffusion Model for Video PredictionZhicheng Zhang, Junyao Hu, Wentao Cheng, Danda Pani Paudel et al.CVPR 2024 · 24 citations
- TEA: Temporal Excitation and Aggregation for Action RecognitionYan Li, Bin Ji, Xintian Shi, Jianguo Zhang et al.CVPR 2020
- Extracting Motion and Appearance via Inter-Frame Attention for Efficient Video Frame InterpolationGuozhen Zhang, Yuhan Zhu, Haonan Wang, Youxin Chen et al.CVPR 2023
- Motion Guided Region Message Passing for Video CaptioningShaoxiang Chen, Yu-Gang JiangICCV 2021 · 71 citations
