VisionLaw: Inferring Interpretable Intrinsic Dynamics from Visual Observations via Bilevel Optimization
Jiajing Lin, Shu Jiang, Qingyuan Zeng, Zhenzhong Wang, Min Jiang
摘要
The intrinsic dynamics of an object governs its physical behavior in the real world, playing a critical role in enabling physically plausible interactive simulation with 3D assets. Existing methods have attempted to infer the intrinsic dynamics of objects from visual observations, but generally face two major challenges: one line of work relies on manually defined constitutive priors, making it difficult to align with actual intrinsic dynamics; the other models intrinsic dynamics using neural networks, resulting in limited interpretability and poor generalization. To address these challenges, we propose VisionLaw, a bilevel optimization framework that infers interpretable expressions of intrinsic dynamics from visual observations. At the upper level, we introduce an LLMs-driven decoupled constitutive evolution strategy, where LLMs are prompted to act as physics experts to generate and revise constitutive laws, with a built-in decoupling mechanism that substantially reduces the search complexity of LLMs. At the lower level, we introduce a vision-guided constitutive evaluation mechanism, which utilizes visual simulation to evaluate the consistency between the generated constitutive law and the underlying intrinsic dynamics, thereby guiding the upper-level evolution. Experiments on both synthetic and real-world datasets demonstrate that VisionLaw can effectively infer interpretable intrinsic dynamics from visual observations. It significantly outperforms existing state-of-the-art methods and exhibits strong generalization for interactive simulation in novel scenarios. Our implementation is available at github.com/JiajingLin/VisionLaw.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- ParticleGS: Learning Neural Gaussian Particle Dynamics from Videos for Prior-free Physical Motion ExtrapolationJinsheng Quan, Qiaowei Miao, Yichao Xu, Zizhuo Lin 等CVPR 2026 · 被引用 5 次
- NeuROK: Generative 4D Neural Object KinematicsChen Geng, Guangzhao He, Yue Gao, Yunzhi Zhang 等CVPR 2026 · 被引用 2 次
它引用的顶会 Paper27
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 被引用 5,687 次
- Instant neural graphics primitives with a multiresolution hash encodingThomas Müller, Alex Evans, Christoph Schied, Alexander KellerSIGGRAPH 2022 · 被引用 4,089 次
- Learning to Simulate Complex Physics with Graph NetworksAlvaro Sanchez-Gonzalez, Jonathan Godwin, Tobias Pfaff, Rex Ying 等ICML 2020 · 被引用 1,439 次
- Learning Mesh-Based Simulation with Graph NetworksTobias Pfaff, Meire Fortunato, Alvaro Sanchez-Gonzalez, Peter W. BattagliaICLR 2021 · 被引用 1,175 次
- DreamGaussian: Generative Gaussian Splatting for Efficient 3D Content CreationJiaxiang Tang, Jiawei Ren, Hang Zhou, Ziwei Liu 等ICLR 2024 · 被引用 955 次
相关 Paper
- Toward Material-Agnostic System Identification From VideosYizhou Zhao, Haoyu Chen, Chunjiang Liu, Zhenyang Li 等ICCV 2025 · 被引用 1 次
- NeuMA: Neural Material Adaptor for Visual Grounding of Intrinsic DynamicsJunyi Cao, Shanyan Guan, Yanhao Ge, Wei Li 等NeurIPS 2024 · 被引用 9 次
- MotionPhysics: Learnable Motion Distillation for Text-Guided SimulationMiaowei Wang, Jakub Zadrozny, Oisin Mac Aodha, Amir VaxmanAAAI 2026
- Visual Grounding of Learned Physical ModelsYunzhu Li, Toru Lin, Kexin Yi, Daniel Bear 等ICML 2020 · 被引用 88 次
- Dynamics: Language-Based Representation for Inferring Rigid-Body Dynamics From VideosChia-Hsiang Kao, Cong Phuoc Huynh, Chien-Yi Wang, Noranart Vesdapunt 等CVPR 2026 · 被引用 2 次
