VGGSfM: Visual Geometry Grounded Deep Structure from Motion
Jianyuan Wang, Nikita Karaev, Christian Rupprecht, David Novotný
Abstract
Structure-from-motion (SfM) is a longstanding problem in the computer vision community, which aims to reconstruct the camera poses and 3D structure of a scene from a set of unconstrained 2D images. Classical frameworks solve this problem in an incremental manner by detecting and matching keypoints, registering images, triangulating 3D points, and conducting bundle adjustment. Recent research efforts have predominantly revolved around harnessing the power of deep learning techniques to enhance specific elements (e.g., keypoint matching), but are still based on the original, non-differentiable pipeline. Instead, we propose a new deep pipeline VGGSfM, where each component is fully differentiable and thus can be trained in an end-to-end manner. To this end, we introduce new mechanisms and simplifications. First, we build on recent advances in deep 2D point tracking to extract reliable pixel-accurate tracks, which eliminates the need for chaining pairwise matches. Furthermore, we recover all cameras simultaneously based on the image and track features instead of gradually registering cameras. Finally, we optimise the cameras and triangulate 3D points via a differentiable bundle adjustment layer. We attain state-of-the-art performance on three popular datasets, CO3D, IMC Phototourism, and ETH3D.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 11df9652-55a2-4565-82d5-6bfa9fea777eCited by top-tier papers94
- Depth Anything 3: Recovering the Visual Space from Any ViewsHaotong Lin, Sili Chen, Jun Hao Liew, Donny Y. Chen et al.ICLR 2026 · 720 citations
- TTT3R: 3D Reconstruction as Test-Time TrainingXingyu Chen, Yue Chen, Yuliang Xiu, Andreas Geiger et al.ICLR 2026 · 139 citations
- MVSplat360: Feed-Forward 360 Scene Synthesis from Sparse ViewsYuedong Chen, Chuanxia Zheng, Haofei Xu, Bohan Zhuang et al.NeurIPS 2024 · 126 citations
- Point3R: Streaming 3D Reconstruction with Explicit Spatial Pointer MemoryYuqi Wu, Wenzhao Zheng, Jie Zhou, Jiwen LuNeurIPS 2025 · 90 citations
- Autoregressive Adversarial Post-Training for Real-Time Interactive Video GenerationShanchuan Lin, Ceyuan Yang, Hao He, Jianwen Jiang et al.NeurIPS 2025 · 89 citations
Builds on23
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- DROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D CamerasZachary Teed, Jia DengNeurIPS 2021 · 1,248 citations
- Common Objects in 3D: Large-Scale Learning and Evaluation of Real-life 3D Category ReconstructionJeremy Reizenstein, Roman Shapovalov, Philipp Henzler, Luca Sbordone et al.ICCV 2021 · 686 citations
- DISK: Learning local features with policy gradientMichal J. Tyszkiewicz, Pascal Fua, Eduard TrullsNeurIPS 2020 · 652 citations
- COTR: Correspondence Transformer for Matching Across ImagesWei Jiang, Eduard Trulls, Jan Hosang, Andrea Tagliasacchi et al.ICCV 2021 · 318 citations
Related papers
- PoseDiffusion: Solving Pose Estimation via Diffusion-aided Bundle AdjustmentJianyuan Wang, Christian Rupprecht, David NovotnýICCV 2023 · 158 citations
- RESfM: Robust Deep Equivariant Structure from MotionFadi Khatib, Yoni Kasten, Dror Moran, Meirav Galun et al.ICLR 2025
- Deep Unsupervised 3D SfM Face Reconstruction Based on Massive Landmark Bundle AdjustmentYuxing Wang, Yawen Lu, Zhihua Xie, Guoyu LuACM MM 2021 · 15 citations
- Deep Permutation Equivariant Structure from MotionDror Moran, Hodaya Koslowsky, Yoni Kasten, Haggai Maron et al.ICCV 2021 · 21 citations
- SpatialTrackerV2: Advancing 3D Point Tracking with Explicit Camera MotionYuxi Xiao, Jianyuan Wang, Nan Xue, Nikita Karaev et al.ICCV 2025 · 6 citations
