3D-IntPhys: Towards More Generalized 3D-grounded Visual Intuitive Physics under Challenging Scenes
Haotian Xue, Antonio Torralba, Josh Tenenbaum, Dan Yamins, Yunzhu Li, Hsiao-Yu Tung
摘要
Given a visual scene, humans have strong intuitions about how a scene can evolve over time under given actions. The intuition, often termed visual intuitive physics, is a critical ability that allows us to make effective plans to manipulate the scene to achieve desired outcomes without relying on extensive trial and error. In this paper, we present a framework capable of learning 3D-grounded visual intuitive physics models from videos of complex scenes with fluids. Our method is composed of a conditional Neural Radiance Field (NeRF)-style visual frontend and a 3D point-based dynamics prediction backend, using which we can impose strong relational and structural inductive bias to capture the structure of the underlying environment. Unlike existing intuitive point-based dynamics works that rely on the supervision of dense point trajectory from simulators, we relax the requirements and only assume access to multi-view RGB images and (imperfect) instance masks acquired using color prior. This enables the proposed model to handle scenarios where accurate point estimation and tracking are hard or impossible. We generate datasets including three challenging scenarios involving fluid, granular materials, and rigid objects in the simulation. The datasets do not include any dense particle information so most previous 3D-based intuitive physics pipelines can barely deal with that. We show our model can make long-horizon future predictions by learning from raw images and significantly outperforms models that do not employ an explicit 3D representation space. We also show that once trained, our model can achieve strong generalization in complex scenarios under extrapolate settings.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- VoMP: Predicting Volumetric Mechanical Property FieldsRishit Dagli, Donglai Xiang, Vismay Modi, Charles Loop 等ICLR 2026 · 被引用 13 次
- Learning 3D-Gaussian Simulators from RGB VideosMikel Zhobro, Andreas René Geist, Georg MartiusICML 2026 · 被引用 8 次
- Learning Physics-Grounded 4D Dynamics with Neural Gaussian Force FieldsShiqian Li, Ruihong Shen, Junfeng Ni, Chang Pan 等ICLR 2026 · 被引用 5 次
- MoSA: Motion-constrained Stress Adaptation for Mitigating Real-to-Sim Gap in Continuum Dynamics via Learning Residual AnisotropyJiaxu Wang, Junhao He, Jingkai SUN, Yi Gu 等ICML 2026
- PIPHEN: Physical Interaction Prediction with Hamiltonian Energy NetworksKewei Chen, Yayu Long, Mingsheng ShangAAAI 2026
它引用的顶会 Paper11
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 被引用 1,852 次
- Learning to Simulate Complex Physics with Graph NetworksAlvaro Sanchez-Gonzalez, Jonathan Godwin, Tobias Pfaff, Rex Ying 等ICML 2020 · 被引用 1,439 次
- Learning Mesh-Based Simulation with Graph NetworksTobias Pfaff, Meire Fortunato, Alvaro Sanchez-Gonzalez, Peter W. BattagliaICLR 2021 · 被引用 1,175 次
- Contrastive Learning of Structured World ModelsThomas N. Kipf, Elise van der Pol, Max WellingICLR 2020 · 被引用 322 次
- Combining Differentiable PDE Solvers and Graph Neural Networks for Fluid Flow PredictionFilipe de Avila Belbute-Peres, Thomas D. Economon, J. Zico KolterICML 2020 · 被引用 271 次
相关 Paper
- Latent Intuitive Physics: Learning to Transfer Hidden Physics from A 3D VideoXiangming Zhu, Huayu Deng, Haochen Yuan, Yunbo Wang 等ICLR 2024 · 被引用 5 次
- Visual Grounding of Learned Physical ModelsYunzhu Li, Toru Lin, Kexin Yi, Daniel Bear 等ICML 2020 · 被引用 88 次
- Dynamic Visual Reasoning by Learning Differentiable Physics Models from Video and LanguageMingyu Ding, Zhenfang Chen, Tao Du, Ping Luo 等NeurIPS 2021 · 被引用 90 次
- Learning 3D Particle-based Simulators from RGB-D VideosWilliam F. Whitney, Tatiana Lopez-Guevara, Tobias Pfaff, Yulia Rubanova 等ICLR 2024 · 被引用 16 次
- Neural Radiance Flow for 4D View Synthesis and Video ProcessingYilun Du, Yinan Zhang, Hong-Xing Yu, Joshua B. Tenenbaum 等ICCV 2021 · 被引用 329 次
