Hierarchical Scene Coordinate Classification and Regression for Visual Localization
Xiaotian Li, Shuzhe Wang, Yi Zhao, Jakob Verbeek, Juho Kannala
Abstract
Visual localization is critical to many applications in computer vision and robotics. To address single-image RGB localization, state-of-the-art feature-based methods match local descriptors between a query image and a pre-built 3D model. Recently, deep neural networks have been exploited to regress the mapping between raw pixels and 3D coordinates in the scene, and thus the matching is implicitly performed by the forward pass through the network. However, in a large and ambiguous environment, learning such a regression task directly can be difficult for a single network. In this work, we present a new hierarchical scene coordinate network to predict pixel scene coordinates in a coarseto-fine manner from a single RGB image. The network consists of a series of output layers, each of them conditioned on the previous ones. The final output layer predicts the 3D coordinates and the others produce progressively finer discrete location labels. The proposed method outperforms the baseline regression-only network and allows us to train compact models which scale robustly to large environments. It sets a new state-of-the-art for single-image RGB localization performance on the 7-Scenes, 12-Scenes, Cambridge Landmarks datasets, and three combined scenes. Moreover, for large-scale outdoor localization on the Aachen Day-Night dataset, we present a hybrid approach which outperforms existing scene coordinate regression methods, and reduces significantly the performance gap w.r.t. explicit feature matching methods. 1 * Work done while JV was at INRIA. 1 Code and materials available at https://aaltovision . github.io/hscnet.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers49
- DUSt3R: Geometric 3D Vision Made EasyShuzhe Wang, Vincent Leroy, Yohann Cabon, Boris Chidlovskii et al.CVPR 2024 · 302 citations
- On the Limits of Pseudo Ground Truth in Visual Camera Re-localisationEric Brachmann, Martin Humenberger, Carsten Rother, Torsten SattlerICCV 2021 · 82 citations
- P2-Net: Joint Description and Detection of Local Features for Pixel and Point MatchingBing Wang, Changhao Chen, Zhaopeng Cui, Jie Qin et al.ICCV 2021 · 75 citations
- PNeRFLoc: Visual Localization with Point-Based Neural Radiance FieldsBoming Zhao, Luwei Yang, Mao Mao, Hujun Bao et al.AAAI 2024 · 31 citations
- Continual Learning for Image-Based Camera LocalizationShuzhe Wang, Zakaria Laskar, Iaroslav Melekhov, Xiaotian Li et al.ICCV 2021 · 31 citations
Builds on6
- Neural-Guided RANSAC: Learning Where to Sample Model HypothesesEric Brachmann, Carsten RotherICCV 2019 · 282 citations
- CamNet: Coarse-to-Fine Retrieval for Camera Re-LocalizationMingyu Ding, Zhe Wang, Jiankai Sun, Jianping Shi et al.ICCV 2019 · 163 citations
- Expert Sample Consensus Applied to Camera Re-LocalizationEric Brachmann, Carsten RotherICCV 2019 · 136 citations
- SANet: Scene Agnostic Network for Camera LocalizationLuwei Yang, Ziqian Bai, Chengzhou Tang, Honghua Li et al.ICCV 2019 · 105 citations
- Fine-Grained Segmentation Networks: Self-Supervised Segmentation for Improved Long-Term Visual LocalizationMåns Larsson, Erik Stenborg, Carl Toft, Lars Hammarstrand et al.ICCV 2019 · 75 citations
Related papers
- Learning Camera Localization via Dense Scene MatchingShitao Tang, Chengzhou Tang, Rui Huang, Siyu Zhu et al.CVPR 2021
- Back to the Feature: Learning Robust Camera Localization From Pixels To PosePaul-Edouard Sarlin, Ajaykumar Unagar, Måns Larsson, Hugo Germain et al.CVPR 2021
- VS-Net: Voting With Segmentation for Visual LocalizationZhaoyang Huang, Han Zhou, Yijin Li, Bangbang Yang et al.CVPR 2021
- UprightNet: Geometry-Aware Camera Orientation Estimation From Single ImagesWenqi Xian, Zhengqi Li, Noah Snavely, Matthew Fisher et al.ICCV 2019 · 52 citations
- GLACE: Global Local Accelerated Coordinate EncodingFangjinhua Wang, Xudong Jiang, Silvano Galliani, Christoph Vogel et al.CVPR 2024 · 24 citations
