Understanding Object Dynamics for Interactive Image-to-Video Synthesis
Andreas Blattmann, Timo Milbich, Michael Dorkenwald, Björn Ommer
Abstract
What would be the effect of locally poking a static scene? We present an approach that learns naturallylooking global articulations caused by a local manipulation at a pixel level. Training requires only videos of moving objects but no information of the underlying manipulation of the physical scene. Our generative model learns to infer natural object dynamics as a response to user interaction and learns about the interrelations between different object body regions. Given a static image of an object and a local poking of a pixel, the approach then predicts how the object would deform over time. In contrast to existing work on video prediction, we do not synthesize arbitrary realistic videos but enable local interactive control of the deformation. Our model is not restricted to particular object categories and can transfer dynamics onto novel unseen object instances. Extensive experiments on diverse objects demonstrate the effectiveness of our approach compared to common video prediction frameworks. Project page is available at https://bit.ly/3cxfA2L .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9e172de0-6158-4594-b91c-0e3316cd61bcCited by top-tier papers18
- Motion-I2V: Consistent and Controllable Image-to-Video Generation with Explicit Motion ModelingXiaoyu Shi, Zhaoyang Huang, Fu-Yun Wang, Weikang Bian et al.SIGGRAPH 2024 · 66 citations
- Make It Move: Controllable Image-to-Video Generation with Text DescriptionsYaosi Hu, Chong Luo, Zhenzhong ChenCVPR 2022 · 56 citations
- iPOKE: Poking a Still Image for Controlled Stochastic Video SynthesisAndreas Blattmann, Timo Milbich, Michael Dorkenwald, Björn OmmerICCV 2021 · 50 citations
- Show Me What and Tell Me How: Video Synthesis via Multimodal ConditioningLigong Han, Jian Ren, Hsin-Ying Lee, Francesco Barbieri et al.CVPR 2022 · 36 citations
- Rethinking Human Evaluation Protocol for Text-to-Video Models: Enhancing Reliability, Reproducibility, and PracticalityTianle Zhang, Langtian Ma, Yuchen Yan, Yuchen Zhang et al.NeurIPS 2024 · 8 citations
Builds on12
- Liquid Warping GAN: A Unified Framework for Human Motion Imitation, Appearance Transfer and Novel View SynthesisWen Liu, Zhixin Piao, Jie Min, Wenhan Luo et al.ICCV 2019 · 285 citations
- Scaling Autoregressive Video ModelsDirk Weissenborn, Oscar Täckström, Jakob UszkoreitICLR 2020 · 252 citations
- Unpaired motion style transfer from video to animationKfir Aberman, Yijia Weng, Dani Lischinski, Daniel Cohen-Or et al.SIGGRAPH 2020 · 178 citations
- Improved Conditional VRNNs for Video PredictionLluís Castrejón, Nicolas Ballas, Aaron C. CourvilleICCV 2019 · 177 citations
- Stochastic Latent Residual Video PredictionJean-Yves Franceschi, Edouard Delasalles, Mickaël Chen, Sylvain Lamprier et al.ICML 2020 · 166 citations
Related papers
- What If: Understanding Motion Through Sparse InteractionsStefan Andreas Baumann, Nick Stracke, Timy Phan, Björn OmmerICCV 2025 · 3 citations
- InterDyn: Controllable Interactive Dynamics with Video Diffusion ModelsRick Akkerman, Haiwen Feng, Michael J. Black, Dimitrios Tzionas et al.CVPR 2025
- MOVES: Manipulated Objects in Video Enable SegmentationRichard E. L. Higgins, David F. FouheyCVPR 2023
- Compositional Video PredictionYufei Ye, Maneesh Singh, Abhinav Gupta, Shubham TulsianiICCV 2019 · 84 citations
- Future Video Synthesis With Object Motion PredictionYue Wu, Rongrong Gao, Jaesik Park, Qifeng ChenCVPR 2020
