Geo4D: Leveraging Video Generators for Geometric 4D Scene Reconstruction
Zeren Jiang, Chuanxia Zheng, Iro Laina, Diane Larlus, Andrea Vedaldi
Abstract
We introduce Geo4D, a method to repurpose video diffusion models for monocular 3D reconstruction of dynamic scenes. By leveraging the strong dynamic priors captured by large-scale pre-trained video models, Geo4D can be trained using only synthetic data while generalizing well to real data in a zero-shot manner. Geo4D predicts several complementary geometric modalities, namely point, disparity, and ray maps. We propose a new multi-modal alignment algorithm to align and fuse these modalities, as well as a sliding window approach at inference time, thus enabling robust and accurate 4D reconstruction of long videos. Extensive experiments across multiple benchmarks show that Geo4D significantly surpasses state-of-the-art video depth estimation methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers23
- Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World ModelingHaoyu Wu, Diankun Wu, Tianyu He, Junliang Guo et al.ICLR 2026 · 89 citations
- STream3R: Scalable Sequential 3D Reconstruction with Causal TransformerYushi Lan, Yihang Luo, Fangzhou Hong, Shangchen Zhou et al.ICLR 2026 · 84 citations
- Efficiently Reconstructing Dynamic Scenes One D4RT at a TimeChuhan Zhang, Guillaume Le Moing, Skanda Koppula, Ignacio Rocco et al.CVPR 2026 · 52 citations
- MoVieS: Motion-Aware 4D Dynamic View Synthesis in One SecondChenguo Lin, Yuchen Lin, Panwang Pan, Yifan Yu et al.CVPR 2026 · 38 citations
- Lyra: Generative 3D Scene Reconstruction via Video Diffusion Model Self-DistillationSherwin Bahmani, Tianchang Shen, Jiawei Ren, Jiahui Huang et al.ICLR 2026 · 33 citations
Builds on63
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan et al.NeurIPS 2022 · 2,948 citations
Related papers
- LaVR: Scene Latent Conditioned Generative Video Trajectory Re-Rendering using Large 4D Reconstruction ModelsMingyang Xie, Numair Khan, Tianfu Wang, Naina Dhingra et al.CVPR 2026 · 10 citations
- C4D: 4D Made from 3D Through Dual CorrespondencesShizun Wang, Zhenxiang Jiang, Xingyi Yang, Xinchao WangICCV 2025 · 4 citations
- MotionCrafter: Dense Geometry and Motion Reconstruction with a 4D VAERuijie Zhu, Jiahao Lu, Wenbo Hu, Xiaoguang Han et al.CVPR 2026 · 3 citations
- GeoMan: Temporally Consistent Human Geometry Estimation Using Image-to-Video DiffusionGwanghyun Kim, Xueting Li, Ye Yuan, Koki Nagano et al.ICCV 2025 · 2 citations
- Geometrycrafter: Consistent Geometry Estimation for Open-World Videos With Diffusion PriorsTian-Xing Xu, Xiangjun Gao, Wenbo Hu, Xiaoyu Li et al.ICCV 2025 · 3 citations
