MMVP: Motion-Matrix-based Video Prediction
Yiqi Zhong, Luming Liang, Ilya Zharkov, Ulrich Neumann
Abstract
A central challenge of video prediction lies where the system has to reason the objects’ future motions from image frames while simultaneously maintaining the consistency of their appearances across frames. This work introduces an end-to-end trainable two-stream video prediction framework, Motion-Matrix-based Video Prediction (MMVP), to tackle this challenge. Unlike previous methods that usually handle motion prediction and appearance maintenance within the same set of modules, MMVP decouples motion and appearance information by constructing appearance-agnostic motion matrices. The motion matrices represent the temporal similarity of each and every pair of feature patches in the input frames, and are the sole input of the motion prediction module in MMVP. This design improves video prediction in both accuracy and efficiency, and reduces the model size. Results of extensive experiments demonstrate that MMVP outperforms state-of-the-art systems on public data sets by non-negligible large margins (≈ 1 db in PSNR, UCF Sports) in significantly smaller model sizes (84% the size or smaller). Please refer to this for the official code and the datasets used in this paper.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2c2ed5c9-2225-4d02-ac25-8312afd0e8e9Cited by top-tier papers7
- STDiff: Spatio-Temporal Diffusion for Continuous Stochastic Video PredictionXi Ye, Guillaume-Alexandre BilodeauAAAI 2024 · 20 citations
- Causal Deciphering and Inpainting in Spatio-Temporal Dynamics via Diffusion ModelYifan Duan, Jian Zhao, pengcheng, Junyuan Mao et al.NeurIPS 2024 · 14 citations
- Motion Graph Unleashed: A Novel Approach to Video PredictionYiqi Zhong, Luming Liang, Bohan Tang, Ilya Zharkov et al.NeurIPS 2024 · 7 citations
- DFDNet: Disentangling and Filtering Dynamics for Enhanced Video PredictionLianqiang Gan, Junyu Lai, Jingze Ju, Lianli Gao et al.AAAI 2025 · 2 citations
- Does Vector Quantization Fail in Spatio-Temporal Forecasting? Exploring a Differentiable Sparse Soft-Vector Quantization ApproachChao Chen, Tian Zhou, Yanjun Zhao, Hui Liu et al.KDD 2025 · 1 citation
Builds on14
- SimVP: Simpler yet Better Video PredictionZhangyang Gao, Cheng Tan, Lirong Wu, Stan Z. LiCVPR 2022 · 313 citations
- Scaling Autoregressive Video ModelsDirk Weissenborn, Oscar Täckström, Jakob UszkoreitICLR 2020 · 252 citations
- Efficient and Information-Preserving Future Frame Prediction and BeyondWei Yu, Yichao Lu, Steve Easterbrook, Sanja FidlerICLR 2020 · 127 citations
- Disentangling Propagation and Generation for Video PredictionHang Gao, Huazhe Xu, Qi-Zhi Cai, Ruth Wang et al.ICCV 2019 · 90 citations
- STRPM: A Spatiotemporal Residual Predictive Model for High-Resolution Video PredictionZheng Chang, Xinfeng Zhang, Shanshe Wang, Siwei Ma et al.CVPR 2022 · 57 citations
Related papers
- Motion-Attentive Transition for Zero-Shot Video Object SegmentationTianfei Zhou, Shunzhou Wang, Yi Zhou, Yazhou Yao et al.AAAI 2020 · 210 citations
- MIMO Is All You Need:A Strong Multi-in-Multi-Out Baseline for Video PredictionShuliang Ning, Mengcheng Lan, Yanran Li, Chaofeng Chen et al.AAAI 2023 · 14 citations
- Video Modeling With Correlation NetworksHeng Wang, Du Tran, Lorenzo Torresani, Matt FeiszliCVPR 2020
- Exploring Spatial-Temporal Multi-Frequency Analysis for High-Fidelity and Temporal-Consistency Video PredictionBeibei Jin, Yu Hu, Qiankun Tang, Jingyu Niu et al.CVPR 2020
- Learning Semantic-Aware Dynamics for Video PredictionXinzhu Bei, Yanchao Yang, Stefano SoattoCVPR 2021
