DynamicVerse: A Physically-Aware Multimodal Framework for 4D World Modeling
Kairun Wen, Yuzhi Huang, Runyu Chen, Hui Zheng, Yunlong Lin, Panwang Pan, Chenxin Li, Wenyan Cong, Jian Jun Zhang, Junbin Lu, Chenguo Lin, Dilin Wang
Abstract
Understanding the dynamic physical world, characterized by its evolving 3D structure, real-world motion, and semantic content with textual descriptions, is crucial for human-agent interaction and enables embodied agents to perceive and act within real environments with human-like capabilities. However, existing datasets are often derived from limited simulators or utilize traditional Structurefrom-Motion for up-to-scale annotation and offer limited descriptive captioning, which restricts the capacity of foundation models to accurately interpret real-world dynamics from monocular videos, commonly sourced from the internet. To bridge these gaps, we introduce DynamicVerse, a physical-scale, multimodal 4D world modeling framework for dynamic real-world video. We employ large vision, geometric, and multimodal models to interpret metric-scale static geometry, real-world dynamic motion, instance-level masks, and holistic descriptive captions. By integrating window-based Bundle Adjustment with global optimization, our method converts long real-world video sequences into a comprehensive 4D multimodal format. DynamicVerse delivers a large-scale dataset consisting of 100K+ videos with 800K+ annotated masks and 10M+ frames from internet videos. Experimental evaluations on three benchmark tasks, namely video depth estimation, camera pose estimation, and camera intrinsics estimation, demonstrate that our 4D modeling achieves superior performance in capturing physical-scale measurements with greater global accuracy than existing methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3ff01442-2af9-49ce-b943-56c9103fe6a3Cited by top-tier papers3
- VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D ReconstructionZhiwen Fan, Jian Zhang, Renjie Li, Junge Zhang et al.CVPR 2026 · 171 citations
- SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial ReasoningJian Zhang, Shijie Zhou, Bangya Liu, Achuta Kadambi et al.CVPR 2026 · 16 citations
- Omni-3DEdit: Generalized Versatile 3D Editing in One-PassLiyi Chen, Pengfei Wang, Guowen Zhang, Zhiyuan Ma et al.CVPR 2026 · 6 citations
Builds on43
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- Depth Anything: Unleashing the Power of Large-Scale Unlabeled DataLihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu et al.CVPR 2024 · 847 citations
- Video Instance SegmentationLinjie Yang, Yuchen Fan, Ning XuICCV 2019 · 615 citations
Related papers
- OmniWorld: A Multi-Domain and Multi-Modal Dataset for 4D World ModelingYang Zhou, Yifan Wang, Jianjun Zhou, Wenzheng Chang et al.ICLR 2026 · 58 citations
- NeoVerse: Enhancing 4D World Model with in-the-wild Monocular VideosYuxue Yang, Lue Fan, Ziqi Shi, Junran Peng et al.CVPR 2026 · 42 citations
- MoSca: Dynamic Gaussian Fusion from Casual Videos via 4D Motion ScaffoldsJiahui Lei, Yijia Weng, Adam W. Harley, Leonidas J. Guibas et al.CVPR 2025
- Uni4D: Unifying Visual Foundation Models for 4D Modeling from a Single VideoDavid Yifan Yao, Albert J. Zhai, Shenlong WangCVPR 2025
- Dynamic Camera Poses and Where to Find ThemChris Rockwell, Joseph Tung, Tsung-Yi Lin, Ming-Yu Liu et al.CVPR 2025
