3D-IntPhys: Towards More Generalized 3D-grounded Visual Intuitive Physics under Challenging Scenes
Haotian Xue, Antonio Torralba, Josh Tenenbaum, Dan Yamins, Yunzhu Li, Hsiao-Yu Tung
Abstract
Given a visual scene, humans have strong intuitions about how a scene can evolve over time under given actions. The intuition, often termed visual intuitive physics, is a critical ability that allows us to make effective plans to manipulate the scene to achieve desired outcomes without relying on extensive trial and error. In this paper, we present a framework capable of learning 3D-grounded visual intuitive physics models from videos of complex scenes with fluids. Our method is composed of a conditional Neural Radiance Field (NeRF)-style visual frontend and a 3D point-based dynamics prediction backend, using which we can impose strong relational and structural inductive bias to capture the structure of the underlying environment. Unlike existing intuitive point-based dynamics works that rely on the supervision of dense point trajectory from simulators, we relax the requirements and only assume access to multi-view RGB images and (imperfect) instance masks acquired using color prior. This enables the proposed model to handle scenarios where accurate point estimation and tracking are hard or impossible. We generate datasets including three challenging scenarios involving fluid, granular materials, and rigid objects in the simulation. The datasets do not include any dense particle information so most previous 3D-based intuitive physics pipelines can barely deal with that. We show our model can make long-horizon future predictions by learning from raw images and significantly outperforms models that do not employ an explicit 3D representation space. We also show that once trained, our model can achieve strong generalization in complex scenarios under extrapolate settings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f520b193-aa5e-49cc-bf60-a2abc8923791Cited by top-tier papers6
- VoMP: Predicting Volumetric Mechanical Property FieldsRishit Dagli, Donglai Xiang, Vismay Modi, Charles Loop et al.ICLR 2026 · 13 citations
- Learning 3D-Gaussian Simulators from RGB VideosMikel Zhobro, Andreas René Geist, Georg MartiusICML 2026 · 8 citations
- Learning Physics-Grounded 4D Dynamics with Neural Gaussian Force FieldsShiqian Li, Ruihong Shen, Junfeng Ni, Chang Pan et al.ICLR 2026 · 5 citations
- MoSA: Motion-constrained Stress Adaptation for Mitigating Real-to-Sim Gap in Continuum Dynamics via Learning Residual AnisotropyJiaxu Wang, Junhao He, Jingkai SUN, Yi Gu et al.ICML 2026
- PIPHEN: Physical Interaction Prediction with Hamiltonian Energy NetworksKewei Chen, Yayu Long, Mingsheng ShangAAAI 2026
Builds on11
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 1,852 citations
- Learning to Simulate Complex Physics with Graph NetworksAlvaro Sanchez-Gonzalez, Jonathan Godwin, Tobias Pfaff, Rex Ying et al.ICML 2020 · 1,439 citations
- Learning Mesh-Based Simulation with Graph NetworksTobias Pfaff, Meire Fortunato, Alvaro Sanchez-Gonzalez, Peter W. BattagliaICLR 2021 · 1,175 citations
- Contrastive Learning of Structured World ModelsThomas N. Kipf, Elise van der Pol, Max WellingICLR 2020 · 322 citations
- Combining Differentiable PDE Solvers and Graph Neural Networks for Fluid Flow PredictionFilipe de Avila Belbute-Peres, Thomas D. Economon, J. Zico KolterICML 2020 · 271 citations
Related papers
- Latent Intuitive Physics: Learning to Transfer Hidden Physics from A 3D VideoXiangming Zhu, Huayu Deng, Haochen Yuan, Yunbo Wang et al.ICLR 2024 · 5 citations
- Visual Grounding of Learned Physical ModelsYunzhu Li, Toru Lin, Kexin Yi, Daniel Bear et al.ICML 2020 · 88 citations
- Dynamic Visual Reasoning by Learning Differentiable Physics Models from Video and LanguageMingyu Ding, Zhenfang Chen, Tao Du, Ping Luo et al.NeurIPS 2021 · 90 citations
- Learning 3D Particle-based Simulators from RGB-D VideosWilliam F. Whitney, Tatiana Lopez-Guevara, Tobias Pfaff, Yulia Rubanova et al.ICLR 2024 · 16 citations
- Neural Radiance Flow for 4D View Synthesis and Video ProcessingYilun Du, Yinan Zhang, Hong-Xing Yu, Joshua B. Tenenbaum et al.ICCV 2021 · 329 citations
