ExtDM: Distribution Extrapolation Diffusion Model for Video Prediction
Zhicheng Zhang, Junyao Hu, Wentao Cheng, Danda Pani Paudel, Jufeng Yang
摘要
Video prediction is a challenging task due to its nature of uncertainty, especially for forecasting a long pe-riod. To model the temporal dynamics, advanced methods benefit from the recent success of diffusion models, and repeatedly refine the predicted future frames with 3D spatiotemporal U-Net. However, there exists a gap between the present and future and the repeated usage of U-Net brings a heavy computation burden. To address this, we propose a diffusion-based video prediction method that predicts future frames by extrapolating the present distribution of features, namely ExtDM. Specifically, our method consists of three components: (i) a motion autoencoder conducts a bijection transformation between video frames and motion cues; (ii) a layered distribution adaptor module extrapolates the present features in the guidance of Gaussian distribution; (iii) a 3D U-Net architecture specialized for jointly fusing guidance and features among the temporal dimension by spatiotemporal-window attention. Extensive experiments on five popular benchmarks covering short- and long-term video prediction verify the effectiveness of ExtDM.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- Video World Models with Long-term Spatial MemoryTong Wu, Shuai Yang, Ryan Po, Yinghao Xu 等NeurIPS 2025 · 被引用 145 次
- Adapt or Perish: Adaptive Sparse Transformer with Attentive Feature Refinement for Image RestorationShihao Zhou, Duosheng Chen, Jinshan Pan, Jinglei Shi 等CVPR 2024 · 被引用 137 次
- FlowDAS: A Stochastic Interpolant-based Framework for Data AssimilationSiyi Chen, Yixuan Jia, Qing Qu, He Sun 等NeurIPS 2025 · 被引用 19 次
- Generative Pre-trained Autoregressive Diffusion TransformerYuan Zhang, Jiacheng Jiang, Guoqing Ma, Zhiying Lu 等NeurIPS 2025 · 被引用 19 次
- Timeline and Boundary Guided Diffusion Network for Video Shadow DetectionHaipeng Zhou, Hongqiu Wang, Tian Ye, Zhaohu Xing 等ACM MM 2024 · 被引用 18 次
它引用的顶会 Paper36
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan 等NeurIPS 2022 · 被引用 2,948 次
- MCVD - Masked Conditional Video Diffusion for Prediction, Generation, and InterpolationVikram Voleti, Alexia Jolicoeur-Martineau, Chris PalNeurIPS 2022 · 被引用 434 次
相关 Paper
- STDiff: Spatio-Temporal Diffusion for Continuous Stochastic Video PredictionXi Ye, Guillaume-Alexandre BilodeauAAAI 2024 · 被引用 20 次
- Grid Diffusion Models for Text-to-Video GenerationTaegyeong Lee, Soyeong Kwon, Taehwan KimCVPR 2024
- Decouple and Track: Benchmarking and Improving Video Diffusion Transformers for Motion TransferQingyu Shi, Jianzong Wu, Jinbin Bai, Jiangning Zhang 等ICCV 2025 · 被引用 1 次
- StoryDiffusion: Consistent Self-Attention for Long-Range Image and Video GenerationYupeng Zhou, Daquan Zhou, Ming-Ming Cheng, Jiashi Feng 等NeurIPS 2024 · 被引用 291 次
- Video Probabilistic Diffusion Models in Projected Latent SpaceSihyun Yu, Kihyuk Sohn, Subin Kim, Jinwoo ShinCVPR 2023
