iPOKE: Poking a Still Image for Controlled Stochastic Video Synthesis
Andreas Blattmann, Timo Milbich, Michael Dorkenwald, Björn Ommer
Abstract
How would a static scene react to a local poke? What are the effects on other parts of an object if you could locally push it? There will be distinctive movement, despite evident variations caused by the stochastic nature of our world. These outcomes are governed by the characteristic kinematics of objects that dictate their overall motion caused by a local interaction. Conversely, the movement of an object provides crucial information about its underlying distinctive kinematics and the interdependencies between its parts. This two-way relation motivates learning a bijective mapping between object kinematics and plausible future image sequences. Therefore, we propose iPOKE – invertible Prediction of Object Kinematics – that, conditioned on an initial frame and a local poke, allows to sample object kinematics and establishes a one-to-one correspondence to the corresponding plausible videos, thereby providing a controlled stochastic video synthesis. In contrast to previous works, we do not generate arbitrary realistic videos, but provide efficient control of movements, while still capturing the stochastic nature of our environment and the diversity of plausible outcomes it entails. Moreover, our approach can transfer kinematics onto novel object instances and is not confined to particular object classes. Our project page is available at https://bit.ly/3dJN4Lf.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9f097ed4-eca4-43cf-ab0d-e2ca0e47d109Cited by top-tier papers21
- Retrieval-Augmented Diffusion ModelsAndreas Blattmann, Robin Rombach, Kaan Oktay, Jonas Müller et al.NeurIPS 2022 · 239 citations
- Disco: Disentangled Control for Realistic Human Dance GenerationTan Wang, Linjie Li, Kevin Lin, Yuanhao Zhai et al.CVPR 2024 · 62 citations
- Automatic Animation of Hair Blowing in Still Portrait PhotosWenpeng Xiao, Wentao Liu, Yitong Wang, Bernard Ghanem et al.ICCV 2023 · 15 citations
- Envisioning the Future, One Step at a TimeStefan Andreas Baumann, Jannik Wiese, Tommaso Martorella, M. Kalayeh et al.CVPR 2026 · 4 citations
- What If: Understanding Motion Through Sparse InteractionsStefan Andreas Baumann, Nick Stracke, Timy Phan, Björn OmmerICCV 2025 · 3 citations
Builds on15
- Scaling Autoregressive Video ModelsDirk Weissenborn, Oscar Täckström, Jakob UszkoreitICLR 2020 · 252 citations
- Unpaired motion style transfer from video to animationKfir Aberman, Yijia Weng, Dani Lischinski, Daniel Cohen-Or et al.SIGGRAPH 2020 · 178 citations
- Improved Conditional VRNNs for Video PredictionLluís Castrejón, Nicolas Ballas, Aaron C. CourvilleICCV 2019 · 177 citations
- Stochastic Latent Residual Video PredictionJean-Yves Franceschi, Edouard Delasalles, Mickaël Chen, Sylvain Lamprier et al.ICML 2020 · 166 citations
- VideoFlow: A Conditional Flow-Based Model for Stochastic Video GenerationManoj Kumar, Mohammad Babaeizadeh, Dumitru Erhan, Chelsea Finn et al.ICLR 2020 · 142 citations
Related papers
- Understanding Object Dynamics for Interactive Image-to-Video SynthesisAndreas Blattmann, Timo Milbich, Michael Dorkenwald, Björn OmmerCVPR 2021
- Stochastic Image-to-Video Synthesis Using cINNsMichael Dorkenwald, Timo Milbich, Andreas Blattmann, Robin Rombach et al.CVPR 2021
- Vid2Game: Controllable Characters Extracted from Real-World VideosOran Gafni, Lior Wolf, Yaniv TaigmanICLR 2020 · 42 citations
- InterDyn: Controllable Interactive Dynamics with Video Diffusion ModelsRick Akkerman, Haiwen Feng, Michael J. Black, Dimitrios Tzionas et al.CVPR 2025
- Learn the Force We Can: Enabling Sparse Motion Control in Multi-Object Video GenerationAram Davtyan, Paolo FavaroAAAI 2024 · 7 citations
