Lune

CVPR2026Top-tier venue

Dynamics: Language-Based Representation for Inferring Rigid-Body Dynamics From Videos

Chia-Hsiang Kao, Cong Phuoc Huynh, Chien-Yi Wang, Noranart Vesdapunt, Stefan Stojanov, Bharath Hariharan, Oleksandr Obiednikov, Ning Zhou

2026Year
2Citations

Abstract

Inferring rigid-body physical states and properties from monocular videos is a fundamental step toward physicsbased perception and simulation. Existing approaches assume specific underlying physical systems, object types, and camera poses, which are unable to generalize to complex real-world settings. We introduce !YNAMICS, a visionlanguage framework that uses language as a unified representation of rigid-body dynamics. Instead of directly predicting parameters, !YNAMICS generates scene configurations in a structured text format for physics simulation. We enhance the model's generalization by integrating natural language motion reasoning and leveraging optical flow as a semantic-agnostic input. On the CLEVRER dataset [59], !YNAMICS achieves a segmentation IoU of 0.30, a 7→ improvement over leading VLMs (InternVL3-8B, Qwen2.5-VL-7B and Claude-4-Sonnet). Further, testtime sampling and evolutionary search further boost performance by 27% and 120% in segmentation IoU, respectively. Finally, we demonstrate strong transfer to a new dataset of 235 real-world rigid-body videos, highlighting the potential of language-driven physics inference for bridging perception and simulation. Additional results and videos are

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 56d9c350-ddf8-435e-a185-1e987ea0cb70

Builds on32

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines