What If: Understanding Motion Through Sparse Interactions
Stefan Andreas Baumann, Nick Stracke, Timy Phan, Björn Ommer
Abstract
Understanding the dynamics of a physical scene involves reasoning about the diverse ways it can potentially change, especially as a result of local interactions. We present the Flow Poke Transformer (FPT), a novel framework for directly predicting the distribution of local motion, conditioned on sparse interactions termed "pokes". Unlike traditional methods that typically only enable dense sampling of a single realization of scene dynamics, FPT provides an interpretable directly accessible representation of multi-modal scene motion, its dependency on physical interactions and the inherent uncertainties of scene dynamics. We also evaluate our model on several downstream tasks to enable comparisons with prior methods and highlight the flexibility of our approach. On dense face motion generation, our generic pre-trained model surpasses specialized baselines. FPT can be fine-tuned in strongly out-of-distribution tasks such as synthetic datasets to enable significant improvements over in-domain methods in articulated object motion estimation. Additionally, predicting explicit motion distributions directly enables our method to achieve competitive performance on tasks like moving part segmentation from pokes which further demonstrates the versatility of our FPT. Code and models are publicly available at https://compvis.github.io/flow-poke-transformer.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6c8b9235-b9a4-4a5c-8dfb-30b7a7b5db57Cited by top-tier papers2
- Envisioning the Future, One Step at a TimeStefan Andreas Baumann, Jannik Wiese, Tommaso Martorella, M. Kalayeh et al.CVPR 2026 · 4 citations
- Physical Object Understanding with a Physically Controllable World ModelRahul Venkatesh, Klemen Kotar, Lilian Naing Chen, Wanhee Lee et al.CVPR 2026 · 1 citation
Builds on25
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
Related papers
- iPOKE: Poking a Still Image for Controlled Stochastic Video SynthesisAndreas Blattmann, Timo Milbich, Michael Dorkenwald, Björn OmmerICCV 2021 · 50 citations
- PhysPT: Physics-aware Pretrained Transformer for Estimating Human Dynamics from Monocular VideosYufei Zhang, Jeffrey O. Kephart, Zijun Cui, Qiang JiCVPR 2024 · 14 citations
- SceneTok: A Compressed, Diffusable Token Space for 3D ScenesMohammad Asim, Christopher Wewer, Jan LenssenCVPR 2026 · 6 citations
- Understanding Object Dynamics for Interactive Image-to-Video SynthesisAndreas Blattmann, Timo Milbich, Michael Dorkenwald, Björn OmmerCVPR 2021
- Uncertainty-Guided Probabilistic Transformer for Complex Action RecognitionHongji Guo, Hanjing Wang, Qiang JiCVPR 2022 · 42 citations
