Dynamics: Language-Based Representation for Inferring Rigid-Body Dynamics From Videos
Chia-Hsiang Kao, Cong Phuoc Huynh, Chien-Yi Wang, Noranart Vesdapunt, Stefan Stojanov, Bharath Hariharan, Oleksandr Obiednikov, Ning Zhou
摘要
Inferring rigid-body physical states and properties from monocular videos is a fundamental step toward physicsbased perception and simulation. Existing approaches assume specific underlying physical systems, object types, and camera poses, which are unable to generalize to complex real-world settings. We introduce !YNAMICS, a visionlanguage framework that uses language as a unified representation of rigid-body dynamics. Instead of directly predicting parameters, !YNAMICS generates scene configurations in a structured text format for physics simulation. We enhance the model's generalization by integrating natural language motion reasoning and leveraging optical flow as a semantic-agnostic input. On the CLEVRER dataset [59], !YNAMICS achieves a segmentation IoU of 0.30, a 7→ improvement over leading VLMs (InternVL3-8B, Qwen2.5-VL-7B and Claude-4-Sonnet). Further, testtime sampling and evolutionary search further boost performance by 27% and 120% in segmentation IoU, respectively. Finally, we demonstrate strong transfer to a new dataset of 235 real-world rigid-body videos, highlighting the potential of language-driven physics inference for bridging perception and simulation. Additional results and videos are
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper32
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun 等ICCV 2021 · 被引用 2,947 次
- CLEVRER: Collision Events for Video Representation and ReasoningKexin Yi, Chuang Gan, Yunzhu Li, Pushmeet Kohli 等ICLR 2020 · 被引用 584 次
- DiffTaichi: Differentiable Programming for Physical SimulationYuanming Hu, Luke Anderson, Tzu-Mao Li, Qi Sun 等ICLR 2020 · 被引用 479 次
- Preference Ranking Optimization for Human AlignmentFeifan Song, Bowen Yu, Minghao Li, Haiyang Yu 等AAAI 2024 · 被引用 357 次
相关 Paper
- VLIPP: Towards Physically Plausible Video Generation with Vision and Language Informed Physical PriorXindi Yang, Baolu Li, Yiming Zhang, Zhenfei Yin 等ICCV 2025 · 被引用 8 次
- Dynamic Visual Reasoning by Learning Differentiable Physics Models from Video and LanguageMingyu Ding, Zhenfang Chen, Tao Du, Ping Luo 等NeurIPS 2021 · 被引用 90 次
- Compositional 4D Dynamic Scenes Understanding with Physics Priors for Video Question AnsweringXingrui Wang, Wufei Ma, Angtian Wang, Shuo Chen 等ICLR 2025
- SIMPACT: Simulation-Enabled Action Planning using Vision-Language ModelsHaowen Liu, Shaoxiong Yao, Haonan Chen, Jiawei Gao 等CVPR 2026 · 被引用 8 次
- Human-Object Interaction via Automatically Designed VLM-Guided Motion PolicyZekai Deng, Ye Shi, Kaiyang Ji, Lan Xu 等ICLR 2026 · 被引用 11 次
