Uni4D: Unifying Visual Foundation Models for 4D Modeling from a Single Video
David Yifan Yao, Albert J. Zhai, Shenlong Wang
2025Year
11Top-tier citations
Abstract
Input Video t Input Video Output 4D Scene Models Output 4D Scene Models Figure 1. Given a casually captured video, Uni4D harnesses pretrained visual foundation models and multi-stage optimization to jointly estimate camera poses, dynamic geometry, and dense 3D motion. The resulting camera poses and geometry are accurate, consistent, and coherent both temporally and spatially.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ff0bdf4a-45a2-4fa3-9e77-81d90f774ad9Cited by top-tier papers11
- MoVieS: Motion-Aware 4D Dynamic View Synthesis in One SecondChenguo Lin, Yuchen Lin, Panwang Pan, Yifan Yu et al.CVPR 2026 · 38 citations
- Vista4D: Video Reshooting with 4D Point CloudsKuan Heng Lin, Zhizheng Liu, Pablo Salamanca, Yash Kant et al.CVPR 2026 · 17 citations
- PAGE-4D: Disentangled Pose and Geometry Estimation for VGGT-4D PerceptionKaichen Zhou, Yuhan Wang, Grace Chen, Gaspard Beaudouin et al.ICLR 2026 · 12 citations
- DynamicVerse: A Physically-Aware Multimodal Framework for 4D World ModelingKairun Wen, Yuzhi Huang, Runyu Chen, Hui Zheng et al.NeurIPS 2025 · 11 citations
- MoRe: Motion-aware Feed-forward 4D Reconstruction TransformerJuntong Fang, Zequn Chen, Weiqi Zhang, Donglin Di et al.CVPR 2026 · 8 citations
Builds on26
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- DROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D CamerasZachary Teed, Jia DengNeurIPS 2021 · 1,248 citations
- Depth Anything: Unleashing the Power of Large-Scale Unlabeled DataLihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu et al.CVPR 2024 · 847 citations
- Dynamic View Synthesis from Dynamic Monocular VideoChen Gao, Ayush Saraf, Johannes Kopf, Jia-Bin HuangICCV 2021 · 522 citations
- 4D Gaussian Splatting for Real-Time Dynamic Scene RenderingGuanjun Wu, Taoran Yi, Jiemin Fang, Lingxi Xie et al.CVPR 2024 · 513 citations
Related papers
- UFO-4D: Unposed Feedforward 4D Reconstruction from Two ImagesJunhwa Hur, Charles Herrmann, Songyou Peng, Philipp Henzler et al.ICLR 2026 · 5 citations
- C4D: 4D Made from 3D Through Dual CorrespondencesShizun Wang, Zhenxiang Jiang, Xingyi Yang, Xinchao WangICCV 2025 · 4 citations
- U4D: Unsupervised 4D Dynamic Scene UnderstandingArmin Mustafa, Chris Russell, Adrian HiltonICCV 2019 · 6 citations
- Joint Optimization for 4D Human-Scene Reconstruction in the WildZhizheng Liu, Joe Lin, Wayne Wu, Bolei ZhouICLR 2026 · 33 citations
- Motion4D: Learning 3D-Consistent Motion and Semantics for 4D Scene UnderstandingHaoran Zhou, Gim Hee LeeNeurIPS 2025 · 3 citations
