Understanding Object Dynamics for Interactive Image-to-Video Synthesis
Andreas Blattmann, Timo Milbich, Michael Dorkenwald, Björn Ommer
摘要
What would be the effect of locally poking a static scene? We present an approach that learns naturallylooking global articulations caused by a local manipulation at a pixel level. Training requires only videos of moving objects but no information of the underlying manipulation of the physical scene. Our generative model learns to infer natural object dynamics as a response to user interaction and learns about the interrelations between different object body regions. Given a static image of an object and a local poking of a pixel, the approach then predicts how the object would deform over time. In contrast to existing work on video prediction, we do not synthesize arbitrary realistic videos but enable local interactive control of the deformation. Our model is not restricted to particular object categories and can transfer dynamics onto novel unseen object instances. Extensive experiments on diverse objects demonstrate the effectiveness of our approach compared to common video prediction frameworks. Project page is available at https://bit.ly/3cxfA2L .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- Motion-I2V: Consistent and Controllable Image-to-Video Generation with Explicit Motion ModelingXiaoyu Shi, Zhaoyang Huang, Fu-Yun Wang, Weikang Bian 等SIGGRAPH 2024 · 被引用 66 次
- Make It Move: Controllable Image-to-Video Generation with Text DescriptionsYaosi Hu, Chong Luo, Zhenzhong ChenCVPR 2022 · 被引用 56 次
- iPOKE: Poking a Still Image for Controlled Stochastic Video SynthesisAndreas Blattmann, Timo Milbich, Michael Dorkenwald, Björn OmmerICCV 2021 · 被引用 50 次
- Show Me What and Tell Me How: Video Synthesis via Multimodal ConditioningLigong Han, Jian Ren, Hsin-Ying Lee, Francesco Barbieri 等CVPR 2022 · 被引用 36 次
- Rethinking Human Evaluation Protocol for Text-to-Video Models: Enhancing Reliability, Reproducibility, and PracticalityTianle Zhang, Langtian Ma, Yuchen Yan, Yuchen Zhang 等NeurIPS 2024 · 被引用 8 次
它引用的顶会 Paper12
- Liquid Warping GAN: A Unified Framework for Human Motion Imitation, Appearance Transfer and Novel View SynthesisWen Liu, Zhixin Piao, Jie Min, Wenhan Luo 等ICCV 2019 · 被引用 285 次
- Scaling Autoregressive Video ModelsDirk Weissenborn, Oscar Täckström, Jakob UszkoreitICLR 2020 · 被引用 252 次
- Unpaired motion style transfer from video to animationKfir Aberman, Yijia Weng, Dani Lischinski, Daniel Cohen-Or 等SIGGRAPH 2020 · 被引用 178 次
- Improved Conditional VRNNs for Video PredictionLluís Castrejón, Nicolas Ballas, Aaron C. CourvilleICCV 2019 · 被引用 177 次
- Stochastic Latent Residual Video PredictionJean-Yves Franceschi, Edouard Delasalles, Mickaël Chen, Sylvain Lamprier 等ICML 2020 · 被引用 166 次
相关 Paper
- What If: Understanding Motion Through Sparse InteractionsStefan Andreas Baumann, Nick Stracke, Timy Phan, Björn OmmerICCV 2025 · 被引用 3 次
- InterDyn: Controllable Interactive Dynamics with Video Diffusion ModelsRick Akkerman, Haiwen Feng, Michael J. Black, Dimitrios Tzionas 等CVPR 2025
- MOVES: Manipulated Objects in Video Enable SegmentationRichard E. L. Higgins, David F. FouheyCVPR 2023
- Compositional Video PredictionYufei Ye, Maneesh Singh, Abhinav Gupta, Shubham TulsianiICCV 2019 · 被引用 84 次
- Future Video Synthesis With Object Motion PredictionYue Wu, Rongrong Gao, Jaesik Park, Qifeng ChenCVPR 2020
