VisionLaw: Inferring Interpretable Intrinsic Dynamics from Visual Observations via Bilevel Optimization
Jiajing Lin, Shu Jiang, Qingyuan Zeng, Zhenzhong Wang, Min Jiang
Abstract
The intrinsic dynamics of an object governs its physical behavior in the real world, playing a critical role in enabling physically plausible interactive simulation with 3D assets. Existing methods have attempted to infer the intrinsic dynamics of objects from visual observations, but generally face two major challenges: one line of work relies on manually defined constitutive priors, making it difficult to align with actual intrinsic dynamics; the other models intrinsic dynamics using neural networks, resulting in limited interpretability and poor generalization. To address these challenges, we propose VisionLaw, a bilevel optimization framework that infers interpretable expressions of intrinsic dynamics from visual observations. At the upper level, we introduce an LLMs-driven decoupled constitutive evolution strategy, where LLMs are prompted to act as physics experts to generate and revise constitutive laws, with a built-in decoupling mechanism that substantially reduces the search complexity of LLMs. At the lower level, we introduce a vision-guided constitutive evaluation mechanism, which utilizes visual simulation to evaluate the consistency between the generated constitutive law and the underlying intrinsic dynamics, thereby guiding the upper-level evolution. Experiments on both synthetic and real-world datasets demonstrate that VisionLaw can effectively infer interpretable intrinsic dynamics from visual observations. It significantly outperforms existing state-of-the-art methods and exhibits strong generalization for interactive simulation in novel scenarios. Our implementation is available at github.com/JiajingLin/VisionLaw.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 98f80402-2c43-4eff-b432-b2b97e01c69cCited by top-tier papers2
- ParticleGS: Learning Neural Gaussian Particle Dynamics from Videos for Prior-free Physical Motion ExtrapolationJinsheng Quan, Qiaowei Miao, Yichao Xu, Zizhuo Lin et al.CVPR 2026 · 5 citations
- NeuROK: Generative 4D Neural Object KinematicsChen Geng, Guangzhao He, Yue Gao, Yunzhi Zhang et al.CVPR 2026 · 2 citations
Builds on27
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- Instant neural graphics primitives with a multiresolution hash encodingThomas Müller, Alex Evans, Christoph Schied, Alexander KellerSIGGRAPH 2022 · 4,089 citations
- Learning to Simulate Complex Physics with Graph NetworksAlvaro Sanchez-Gonzalez, Jonathan Godwin, Tobias Pfaff, Rex Ying et al.ICML 2020 · 1,439 citations
- Learning Mesh-Based Simulation with Graph NetworksTobias Pfaff, Meire Fortunato, Alvaro Sanchez-Gonzalez, Peter W. BattagliaICLR 2021 · 1,175 citations
- DreamGaussian: Generative Gaussian Splatting for Efficient 3D Content CreationJiaxiang Tang, Jiawei Ren, Hang Zhou, Ziwei Liu et al.ICLR 2024 · 955 citations
Related papers
- Toward Material-Agnostic System Identification From VideosYizhou Zhao, Haoyu Chen, Chunjiang Liu, Zhenyang Li et al.ICCV 2025 · 1 citation
- NeuMA: Neural Material Adaptor for Visual Grounding of Intrinsic DynamicsJunyi Cao, Shanyan Guan, Yanhao Ge, Wei Li et al.NeurIPS 2024 · 9 citations
- MotionPhysics: Learnable Motion Distillation for Text-Guided SimulationMiaowei Wang, Jakub Zadrozny, Oisin Mac Aodha, Amir VaxmanAAAI 2026
- Visual Grounding of Learned Physical ModelsYunzhu Li, Toru Lin, Kexin Yi, Daniel Bear et al.ICML 2020 · 88 citations
- Dynamics: Language-Based Representation for Inferring Rigid-Body Dynamics From VideosChia-Hsiang Kao, Cong Phuoc Huynh, Chien-Yi Wang, Noranart Vesdapunt et al.CVPR 2026 · 2 citations
