Geometrycrafter: Consistent Geometry Estimation for Open-World Videos With Diffusion Priors
Tian-Xing Xu, Xiangjun Gao, Wenbo Hu, Xiaoyu Li, Song-Hai Zhang, Ying Shan
Abstract
Despite remarkable advancements in video depth estimation, existing methods exhibit inherent limitations in achieving geometric fidelity through the affine-invariant predictions, limiting their applicability in reconstruction and other metrically grounded downstream tasks. We propose GeometryCrafter, a novel framework that recovers high-fidelity point map sequences with temporal coherence from open-world videos, enabling accurate 3D/4D reconstruction, camera parameter estimation, and other depth-based applications. At the core of our approach lies a point map Variational Autoencoder (VAE) that learns a latent space agnostic to video latent distributions for effective point map encoding and decoding. Leveraging the VAE, we train a video diffusion model to model the distribution of point map sequences conditioned on the input videos. Extensive evaluations on diverse datasets demonstrate that GeometryCrafter achieves state-of-the-art 3D accuracy, temporal consistency, and generalization capability.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 26f310a0-5d0f-4c7e-b2af-0e574b327affCited by top-tier papers16
- Pixel-Perfect Depth with Semantics-Prompted Diffusion TransformersGangwei Xu, Haotong Lin, Hongcheng Luo, Xianqi Wang et al.NeurIPS 2025 · 58 citations
- MoVieS: Motion-Aware 4D Dynamic View Synthesis in One SecondChenguo Lin, Yuchen Lin, Panwang Pan, Yifan Yu et al.CVPR 2026 · 38 citations
- Vista4D: Video Reshooting with 4D Point CloudsKuan Heng Lin, Zhizheng Liu, Pablo Salamanca, Yash Kant et al.CVPR 2026 · 17 citations
- Mesh4D: 4D Mesh Reconstruction and Tracking from Monocular VideoZeren Jiang, Chuanxia Zheng, Iro Laina, Diane Larlus et al.CVPR 2026 · 13 citations
- PAGE-4D: Disentangled Pose and Geometry Estimation for VGGT-4D PerceptionKaichen Zhou, Yuhan Wang, Grace Chen, Gaspard Beaudouin et al.ICLR 2026 · 12 citations
Builds on40
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- Elucidating the Design Space of Diffusion-Based Generative ModelsTero Karras, Miika Aittala, Timo Aila, Samuli LaineNeurIPS 2022 · 3,959 citations
Related papers
- MotionCrafter: Dense Geometry and Motion Reconstruction with a 4D VAERuijie Zhu, Jiahao Lu, Wenbo Hu, Xiaoguang Han et al.CVPR 2026 · 3 citations
- DepthCrafter: Generating Consistent Long Depth Sequences for Open-world VideosWenbo Hu, Xiangjun Gao, Xiaoyu Li, Sijie Zhao et al.CVPR 2025
- NormalCrafter: Learning Temporally Consistent Normals from Video Diffusion PriorsYanrui Bin, Wenbo Hu, Haoyuan Wang, Xinya Chen et al.ICCV 2025 · 27 citations
- VerseCrafter: Dynamic Realistic Video World Model with 4D Geometric ControlSixiao Zheng, Minghao Yin, Wenbo Hu, Xiaoyu Li et al.CVPR 2026 · 27 citations
- Geo4D: Leveraging Video Generators for Geometric 4D Scene ReconstructionZeren Jiang, Chuanxia Zheng, Iro Laina, Diane Larlus et al.ICCV 2025 · 9 citations
