Structured Object-Aware Physics Prediction for Video Modeling and Planning
Jannik Kossen, Karl Stelzner, Marcel Hussing, Claas Voelcker, Kristian Kersting
Abstract
When humans observe a physical system, they can easily locate objects, understand their interactions, and anticipate future behavior, even in settings with complicated and previously unseen interactions. For computers, however, learning such models from videos in an unsupervised fashion is an unsolved research problem. In this paper, we present STOVE, a novel state-space model for videos, which explicitly reasons about objects and their positions, velocities, and interactions. It is constructed by combining an image model and a dynamics model in compositional manner and improves on previous work by reusing the dynamics model for inference, accelerating and regularizing training. STOVE predicts videos with convincing physical behavior over hundreds of timesteps, outperforms previous unsupervised models, and even approaches the performance of supervised baselines. We further demonstrate the strength of our model as a simulator for sample efficient model-based control in a task with heavily interacting objects.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers23
- Simple Unsupervised Object-Centric Learning for Complex and Naturalistic VideosGautam Singh, Yi-Fu Wu, Sungjin AhnNeurIPS 2022 · 182 citations
- SlotDiffusion: Object-Centric Generative Modeling with Diffusion ModelsZiyi Wu, Jingyu Hu, Wuyue Lu, Igor Gilitschenski et al.NeurIPS 2023 · 106 citations
- Improving Generative Imagination in Object-Centric World ModelsZhixuan Lin, Yi-Fu Wu, Skand Vishwanath Peri, Bofeng Fu et al.ICML 2020 · 97 citations
- Generalization and Robustness Implications in Object-Centric LearningAndrea Dittadi, Samuele S. Papa, Michele De Vita, Bernhard Schölkopf et al.ICML 2022 · 87 citations
- RELATE: Physically Plausible Multi-Object Scene Synthesis Using Structured Latent SpacesSébastien Ehrhardt, Oliver Groth, Áron Monszpart, Martin Engelcke et al.NeurIPS 2020 · 61 citations
Builds on2
- Model Based Reinforcement Learning for AtariLukasz Kaiser, Mohammad Babaeizadeh, Piotr Milos, Blazej Osinski et al.ICLR 2020 · 969 citations
- GENESIS: Generative Scene Inference and Sampling with Object-Centric Latent RepresentationsMartin Engelcke, Adam R. Kosiorek, Oiwi Parker Jones, Ingmar PosnerICLR 2020 · 334 citations
Related papers
- PlaySlot: Learning Inverse Latent Dynamics for Controllable Object-Centric Video Prediction and PlanningAngel Villar-Corrales, Sven BehnkeICML 2025
- Physics-as-Inverse-Graphics: Unsupervised Physical Parameter Estimation from VideoMiguel Jaques, Michael Burke, Timothy M. HospedalesICLR 2020 · 58 citations
- Hierarchical Relational InferenceAleksandar Stanic, Sjoerd van Steenkiste, Jürgen SchmidhuberAAAI 2021 · 17 citations
- Videos are Sample-Efficient Supervisions: Behavior Cloning from Videos via Latent RepresentationsXin Liu, Haoran Li, Dongbin ZhaoNeurIPS 2025 · 5 citations
- SIMONe: View-Invariant, Temporally-Abstracted Object Representations via Unsupervised Video DecompositionRishabh Kabra, Daniel Zoran, Goker Erdogan, Loic Matthey et al.NeurIPS 2021 · 90 citations
