Visual Grounding of Learned Physical Models
Yunzhu Li, Toru Lin, Kexin Yi, Daniel Bear, Daniel Yamins, Jiajun Wu, Joshua B. Tenenbaum, Antonio Torralba
Abstract
Humans intuitively recognize objects' physical properties and predict their motion, even when the objects are engaged in complicated interactions. The abilities to perform physical reasoning and to adapt to new environments, while intrinsic to humans, remain challenging to state-of-the-art computational models. In this work, we present a neural model that simultaneously reasons about physics and makes future predictions based on visual and dynamics priors. The visual prior predicts a particle-based representation of the system from visual observations. An inference module operates on those particles, predicting and refining estimates of particle locations, object states, and physical parameters, subject to the constraints imposed by the dynamics prior, which we refer to as visual grounding. We demonstrate the effectiveness of our method in environments involving rigid objects, deformable materials, and fluids. Experiments show that our model can infer the physical properties within a few observations, which allows the model to quickly adapt to unseen scenarios and make accurate predictions into the future.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9b8abe5a-6142-40e3-911a-895659ddf0f1Cited by top-tier papers21
- Causal Discovery in Physical Systems from VideosYunzhu Li, Antonio Torralba, Anima Anandkumar, Dieter Fox et al.NeurIPS 2020 · 133 citations
- gradSim: Differentiable simulation for system identification and visuomotor controlJ. Krishna Murthy, Miles Macklin, Florian Golemo, Vikram Voleti et al.ICLR 2021 · 130 citations
- Dynamic Visual Reasoning by Learning Differentiable Physics Models from Video and LanguageMingyu Ding, Zhenfang Chen, Tao Du, Ping Luo et al.NeurIPS 2021 · 90 citations
- Learning Physical Graph Representations from Visual ScenesDaniel Bear, Chaofei Fan, Damian Mrowca, Yunzhu Li et al.NeurIPS 2020 · 88 citations
- Learning Physical Dynamics with Subequivariant Graph Neural NetworksJiaqi Han, Wenbing Huang, Hengbo Ma, Jiachen Li et al.NeurIPS 2022 · 72 citations
Builds on4
- Learning to Simulate Complex Physics with Graph NetworksAlvaro Sanchez-Gonzalez, Jonathan Godwin, Tobias Pfaff, Rex Ying et al.ICML 2020 · 1,439 citations
- CLEVRER: Collision Events for Video Representation and ReasoningKexin Yi, Chuang Gan, Yunzhu Li, Pushmeet Kohli et al.ICLR 2020 · 584 citations
- Lagrangian Fluid Simulation with Continuous ConvolutionsBenjamin Ummenhofer, Lukas Prantl, Nils Thuerey, Vladlen KoltunICLR 2020 · 211 citations
- Learning Compositional Koopman Operators for Model-Based ControlYunzhu Li, Hao He, Jiajun Wu, Dina Katabi et al.ICLR 2020 · 135 citations
Related papers
- 3D-IntPhys: Towards More Generalized 3D-grounded Visual Intuitive Physics under Challenging ScenesHaotian Xue, Antonio Torralba, Josh Tenenbaum, Dan Yamins et al.NeurIPS 2023 · 19 citations
- SlotPi: Physics-informed Object-centric Reasoning ModelsJian Li, Han Wan, Ning Lin, Yu-Liang Zhan et al.KDD 2025
- NeuroFluid: Fluid Dynamics Grounding with Particle-Driven Neural Radiance FieldsShanyan Guan, Huayu Deng, Yunbo Wang, Xiaokang YangICML 2022 · 51 citations
- Physics-as-Inverse-Graphics: Unsupervised Physical Parameter Estimation from VideoMiguel Jaques, Michael Burke, Timothy M. HospedalesICLR 2020 · 58 citations
- Hierarchical Relational InferenceAleksandar Stanic, Sjoerd van Steenkiste, Jürgen SchmidhuberAAAI 2021 · 17 citations
