Stochastic Image-to-Video Synthesis Using cINNs
Michael Dorkenwald, Timo Milbich, Andreas Blattmann, Robin Rombach, Konstantinos G. Derpanis, Björn Ommer
Abstract
Video understanding calls for a model to learn the characteristic interplay between static scene content and its dynamics: Given an image, the model must be able to predict a future progression of the portrayed scene and, conversely, a video should be explained in terms of its static image content and all the remaining characteristics not present in the initial frame. This naturally suggests a bijective mapping between the video domain and the static content as well as residual information. In contrast to common stochastic image-to-video synthesis, such a model does not merely generate arbitrary videos progressing the initial image. Given this image, it rather provides a one-to-one mapping between the residual vectors and the video with stochastic outcomes when sampling. The approach is naturally implemented using a conditional invertible neural network (cINN) that can explain videos by independently modelling static and other video characteristics, thus laying the basis for controlled video synthesis. Experiments on diverse video datasets demonstrate the effectiveness of our approach in terms of both the quality and diversity of the synthesized results. Our project page is available at https://bit.ly/3dg90fV . * Indicates equal supervision.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4a706106-7ed6-4dfc-989c-5ba61678cebeCited by top-tier papers19
- NExT-GPT: Any-to-Any Multimodal LLMShengqiong Wu, Hao Fei, Leigang Qu, Wei Ji et al.ICML 2024 · 786 citations
- Make It Move: Controllable Image-to-Video Generation with Text DescriptionsYaosi Hu, Chong Luo, Zhenzhong ChenCVPR 2022 · 56 citations
- iPOKE: Poking a Still Image for Controlled Stochastic Video SynthesisAndreas Blattmann, Timo Milbich, Michael Dorkenwald, Björn OmmerICCV 2021 · 50 citations
- MaskViT: Masked Visual Pre-Training for Video PredictionAgrim Gupta, Stephen Tian, Yunzhi Zhang, Jiajun Wu et al.ICLR 2023 · 45 citations
- Show Me What and Tell Me How: Video Synthesis via Multimodal ConditioningLigong Han, Jian Ren, Hsin-Ying Lee, Francesco Barbieri et al.CVPR 2022 · 36 citations
Builds on13
- Liquid Warping GAN: A Unified Framework for Human Motion Imitation, Appearance Transfer and Novel View SynthesisWen Liu, Zhixin Piao, Jie Min, Wenhan Luo et al.ICCV 2019 · 285 citations
- Scaling Autoregressive Video ModelsDirk Weissenborn, Oscar Täckström, Jakob UszkoreitICLR 2020 · 252 citations
- Drive&Act: A Multi-Modal Dataset for Fine-Grained Driver Behavior Recognition in Autonomous VehiclesManuel Martin, Alina Roitberg, Monica Haurilet, Matthias Horne et al.ICCV 2019 · 235 citations
- Improved Conditional VRNNs for Video PredictionLluís Castrejón, Nicolas Ballas, Aaron C. CourvilleICCV 2019 · 177 citations
- Stochastic Latent Residual Video PredictionJean-Yves Franceschi, Edouard Delasalles, Mickaël Chen, Sylvain Lamprier et al.ICML 2020 · 166 citations
Related papers
- STDiff: Spatio-Temporal Diffusion for Continuous Stochastic Video PredictionXi Ye, Guillaume-Alexandre BilodeauAAAI 2024 · 20 citations
- Decouple Content and Motion for Conditional Image-to-Video GenerationCuifeng Shen, Yulu Gan, Chen Chen, Xiongwei Zhu et al.AAAI 2024 · 13 citations
- OmniVDiff: Omni Controllable Video Diffusion for Generation and UnderstandingDianbing Xi, Jiepeng Wang, Yuanzhi Liang, Xi Qiu et al.AAAI 2026 · 14 citations
- CCVS: Context-aware Controllable Video SynthesisGuillaume Le Moing, Jean Ponce, Cordelia SchmidNeurIPS 2021 · 98 citations
- Behavior-Driven Synthesis of Human DynamicsAndreas Blattmann, Timo Milbich, Michael Dorkenwald, Björn OmmerCVPR 2021
