Pixel-Perfect Structure-from-Motion with Featuremetric Refinement
Philipp Lindenberger, Paul-Edouard Sarlin, Viktor Larsson, Marc Pollefeys
Abstract
Finding local features that are repeatable across multiple views is a cornerstone of sparse 3D reconstruction. The classical image matching paradigm detects keypoints per-image once and for all, which can yield poorly-localized features and propagate large errors to the final geometry. In this paper, we refine two key steps of structure-from-motion by a direct alignment of low-level image information from multiple views: we first adjust the initial keypoint locations prior to any geometric estimation, and subsequently refine points and camera poses as a post-processing. This refinement is robust to large detection noise and appearance changes, as it optimizes a featuremetric error based on dense features predicted by a neural network. This significantly improves the accuracy of camera poses and scene geometry for a wide range of keypoint detectors, challenging viewing conditions, and off-the-shelf deep features. Our system easily scales to large image collections, enabling pixel-perfect crowd-sourced localization at scale. Our code is publicly available at github.com/cvg/pixel-perfect-sfm as an add-on to the popular SfM software COLMAP.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers84
- LightGlue: Local Feature Matching at Light SpeedPhilipp Lindenberger, Paul-Edouard Sarlin, Marc PollefeysICCV 2023 · 936 citations
- Mega-NeRF: Scalable Construction of Large-Scale NeRFs for Virtual Fly- ThroughsHaithem Turki, Deva Ramanan, Mahadev SatyanarayananCVPR 2022 · 364 citations
- Deep Patch Visual OdometryZachary Teed, Lahav Lipson, Jia DengNeurIPS 2023 · 323 citations
- DUSt3R: Geometric 3D Vision Made EasyShuzhe Wang, Vincent Leroy, Yohann Cabon, Boris Chidlovskii et al.CVPR 2024 · 302 citations
- OnePose++: Keypoint-Free One-Shot Object Pose Estimation without CAD ModelsXingyi He, Jiaming Sun, Yuang Wang, Di Huang et al.NeurIPS 2022 · 190 citations
Builds on9
- DISK: Learning local features with policy gradientMichal J. Tyszkiewicz, Pascal Fua, Eduard TrullsNeurIPS 2020 · 652 citations
- Dual-Resolution Correspondence NetworksXinghui Li, Kai Han, Shuda Li, Victor PrisacariuNeurIPS 2020 · 207 citations
- Neural Reprojection Error: Merging Feature Learning and Camera Pose EstimationHugo Germain, Vincent Lepetit, Guillaume BourmaudCVPR 2021
- SuperGlue: Learning Feature Matching With Graph Neural NetworksPaul-Edouard Sarlin, Daniel DeTone, Tomasz Malisiewicz, Andrew RabinovichCVPR 2020
- Patch2Pix: Epipolar-Guided Pixel-Level CorrespondencesQunjie Zhou, Torsten Sattler, Laura Leal-TaixéCVPR 2021
Related papers
- Detector-Free Structure from MotionXingyi He, Jiaming Sun, Yifan Wang, Sida Peng et al.CVPR 2024
- Back to the Feature: Learning Robust Camera Localization From Pixels To PosePaul-Edouard Sarlin, Ajaykumar Unagar, Måns Larsson, Hugo Germain et al.CVPR 2021
- Rooms from Motion: Un-posed Indoor 3D Object Detection as Localization and MappingJustin Lazarow, Kai Kang, Afshin DehghanNeurIPS 2025 · 2 citations
- Sparse-View Localization via Online Neural 3D RegressionLudvig Dillén, Magnus Oskarsson, Viktor LarssonCVPR 2026
- VGGSfM: Visual Geometry Grounded Deep Structure from MotionJianyuan Wang, Nikita Karaev, Christian Rupprecht, David NovotnýCVPR 2024 · 48 citations
