STDiff: Spatio-Temporal Diffusion for Continuous Stochastic Video Prediction
Xi Ye, Guillaume-Alexandre Bilodeau
摘要
Predicting future frames of a video is challenging because it is difficult to learn the uncertainty of the underlying factors influencing their contents. In this paper, we propose a novel video prediction model, which has infinite-dimensional latent variables over the spatio-temporal domain. Specifically, we first decompose the video motion and content information, then take a neural stochastic differential equation to predict the temporal motion information, and finally, an image diffusion model autoregressively generates the video frame by conditioning on the predicted motion feature and the previous frame. The better expressiveness and stronger stochasticity learning capability of our model lead to state-of-theart video prediction performances. As well, our model is able to achieve temporal continuous prediction, i.e., predicting in an unsupervised way the future video frames with an arbitrarily high frame rate. Our code is available at https: //github.com/XiYe20/STDiffProject .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- GenRec: Unifying Video Generation and Recognition with Diffusion ModelsZejia Weng, Xitong Yang, Zhen Xing, Zuxuan Wu 等NeurIPS 2024 · 被引用 19 次
- Causal-Entity Reflected Egocentric Traffic Accident Video SynthesisLei-Lei Li, Jianwu Fang, Junbin Xiao, Shanmin Pang 等ICCV 2025 · 被引用 4 次
- DFDNet: Disentangling and Filtering Dynamics for Enhanced Video PredictionLianqiang Gan, Junyu Lai, Jingze Ju, Lianli Gao 等AAAI 2025 · 被引用 2 次
- CRONOS: Continuous time reconstruction for 4D medical longitudinal seriesNico Disch, Saikat Roy, Constantin Ulrich, Yannick Kirchhoff 等ICLR 2026 · 被引用 2 次
- SyncVP: Joint Diffusion for Synchronous Multi-Modal Video PredictionEnrico Pallotta, Sina Mokhtarzadeh Azar, Shuai Li, Olga Zatsarynna 等CVPR 2025
它引用的顶会 Paper17
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar 等ICLR 2021 · 被引用 1,270 次
- Autoregressive Denoising Diffusion Models for Multivariate Probabilistic Time Series ForecastingKashif Rasul, Calvin Seward, Ingmar Schuster, Roland VollgrafICML 2021 · 被引用 500 次
- Flexible Diffusion Modeling of Long VideosWilliam Harvey, Saeid Naderiparizi, Vaden Masrani, Christian Weilbach 等NeurIPS 2022 · 被引用 384 次
- A Variational Perspective on Diffusion-Based Generative Models and Score MatchingChin-Wei Huang, Jae Hyun Lim, Aaron C. CourvilleNeurIPS 2021 · 被引用 246 次
相关 Paper
- Stochastic Latent Residual Video PredictionJean-Yves Franceschi, Edouard Delasalles, Mickaël Chen, Sylvain Lamprier 等ICML 2020 · 被引用 166 次
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan 等NeurIPS 2022 · 被引用 2,948 次
- ExtDM: Distribution Extrapolation Diffusion Model for Video PredictionZhicheng Zhang, Junyao Hu, Wentao Cheng, Danda Pani Paudel 等CVPR 2024 · 被引用 24 次
- STDD: Spatio-Temporal Dual Diffusion for Video GenerationShuaizhen Yao, Xiaoya Zhang, Xin Liu, Mengyi Liu 等CVPR 2025
- VideoFlow: A Conditional Flow-Based Model for Stochastic Video GenerationManoj Kumar, Mohammad Babaeizadeh, Dumitru Erhan, Chelsea Finn 等ICLR 2020 · 被引用 142 次
