MMVP: Motion-Matrix-based Video Prediction
Yiqi Zhong, Luming Liang, Ilya Zharkov, Ulrich Neumann
摘要
A central challenge of video prediction lies where the system has to reason the objects’ future motions from image frames while simultaneously maintaining the consistency of their appearances across frames. This work introduces an end-to-end trainable two-stream video prediction framework, Motion-Matrix-based Video Prediction (MMVP), to tackle this challenge. Unlike previous methods that usually handle motion prediction and appearance maintenance within the same set of modules, MMVP decouples motion and appearance information by constructing appearance-agnostic motion matrices. The motion matrices represent the temporal similarity of each and every pair of feature patches in the input frames, and are the sole input of the motion prediction module in MMVP. This design improves video prediction in both accuracy and efficiency, and reduces the model size. Results of extensive experiments demonstrate that MMVP outperforms state-of-the-art systems on public data sets by non-negligible large margins (≈ 1 db in PSNR, UCF Sports) in significantly smaller model sizes (84% the size or smaller). Please refer to this for the official code and the datasets used in this paper.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- STDiff: Spatio-Temporal Diffusion for Continuous Stochastic Video PredictionXi Ye, Guillaume-Alexandre BilodeauAAAI 2024 · 被引用 20 次
- Causal Deciphering and Inpainting in Spatio-Temporal Dynamics via Diffusion ModelYifan Duan, Jian Zhao, pengcheng, Junyuan Mao 等NeurIPS 2024 · 被引用 14 次
- Motion Graph Unleashed: A Novel Approach to Video PredictionYiqi Zhong, Luming Liang, Bohan Tang, Ilya Zharkov 等NeurIPS 2024 · 被引用 7 次
- DFDNet: Disentangling and Filtering Dynamics for Enhanced Video PredictionLianqiang Gan, Junyu Lai, Jingze Ju, Lianli Gao 等AAAI 2025 · 被引用 2 次
- Does Vector Quantization Fail in Spatio-Temporal Forecasting? Exploring a Differentiable Sparse Soft-Vector Quantization ApproachChao Chen, Tian Zhou, Yanjun Zhao, Hui Liu 等KDD 2025 · 被引用 1 次
它引用的顶会 Paper14
- SimVP: Simpler yet Better Video PredictionZhangyang Gao, Cheng Tan, Lirong Wu, Stan Z. LiCVPR 2022 · 被引用 313 次
- Scaling Autoregressive Video ModelsDirk Weissenborn, Oscar Täckström, Jakob UszkoreitICLR 2020 · 被引用 252 次
- Efficient and Information-Preserving Future Frame Prediction and BeyondWei Yu, Yichao Lu, Steve Easterbrook, Sanja FidlerICLR 2020 · 被引用 127 次
- Disentangling Propagation and Generation for Video PredictionHang Gao, Huazhe Xu, Qi-Zhi Cai, Ruth Wang 等ICCV 2019 · 被引用 90 次
- STRPM: A Spatiotemporal Residual Predictive Model for High-Resolution Video PredictionZheng Chang, Xinfeng Zhang, Shanshe Wang, Siwei Ma 等CVPR 2022 · 被引用 57 次
相关 Paper
- Motion-Attentive Transition for Zero-Shot Video Object SegmentationTianfei Zhou, Shunzhou Wang, Yi Zhou, Yazhou Yao 等AAAI 2020 · 被引用 210 次
- MIMO Is All You Need:A Strong Multi-in-Multi-Out Baseline for Video PredictionShuliang Ning, Mengcheng Lan, Yanran Li, Chaofeng Chen 等AAAI 2023 · 被引用 14 次
- Video Modeling With Correlation NetworksHeng Wang, Du Tran, Lorenzo Torresani, Matt FeiszliCVPR 2020
- Exploring Spatial-Temporal Multi-Frequency Analysis for High-Fidelity and Temporal-Consistency Video PredictionBeibei Jin, Yu Hu, Qiankun Tang, Jingyu Niu 等CVPR 2020
- Learning Semantic-Aware Dynamics for Video PredictionXinzhu Bei, Yanchao Yang, Stefano SoattoCVPR 2021
