Hierarchical Scene Coordinate Classification and Regression for Visual Localization
Xiaotian Li, Shuzhe Wang, Yi Zhao, Jakob Verbeek, Juho Kannala
摘要
Visual localization is critical to many applications in computer vision and robotics. To address single-image RGB localization, state-of-the-art feature-based methods match local descriptors between a query image and a pre-built 3D model. Recently, deep neural networks have been exploited to regress the mapping between raw pixels and 3D coordinates in the scene, and thus the matching is implicitly performed by the forward pass through the network. However, in a large and ambiguous environment, learning such a regression task directly can be difficult for a single network. In this work, we present a new hierarchical scene coordinate network to predict pixel scene coordinates in a coarseto-fine manner from a single RGB image. The network consists of a series of output layers, each of them conditioned on the previous ones. The final output layer predicts the 3D coordinates and the others produce progressively finer discrete location labels. The proposed method outperforms the baseline regression-only network and allows us to train compact models which scale robustly to large environments. It sets a new state-of-the-art for single-image RGB localization performance on the 7-Scenes, 12-Scenes, Cambridge Landmarks datasets, and three combined scenes. Moreover, for large-scale outdoor localization on the Aachen Day-Night dataset, we present a hybrid approach which outperforms existing scene coordinate regression methods, and reduces significantly the performance gap w.r.t. explicit feature matching methods. 1 * Work done while JV was at INRIA. 1 Code and materials available at https://aaltovision . github.io/hscnet.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper49
- DUSt3R: Geometric 3D Vision Made EasyShuzhe Wang, Vincent Leroy, Yohann Cabon, Boris Chidlovskii 等CVPR 2024 · 被引用 302 次
- On the Limits of Pseudo Ground Truth in Visual Camera Re-localisationEric Brachmann, Martin Humenberger, Carsten Rother, Torsten SattlerICCV 2021 · 被引用 82 次
- P2-Net: Joint Description and Detection of Local Features for Pixel and Point MatchingBing Wang, Changhao Chen, Zhaopeng Cui, Jie Qin 等ICCV 2021 · 被引用 75 次
- PNeRFLoc: Visual Localization with Point-Based Neural Radiance FieldsBoming Zhao, Luwei Yang, Mao Mao, Hujun Bao 等AAAI 2024 · 被引用 31 次
- Continual Learning for Image-Based Camera LocalizationShuzhe Wang, Zakaria Laskar, Iaroslav Melekhov, Xiaotian Li 等ICCV 2021 · 被引用 31 次
它引用的顶会 Paper6
- Neural-Guided RANSAC: Learning Where to Sample Model HypothesesEric Brachmann, Carsten RotherICCV 2019 · 被引用 282 次
- CamNet: Coarse-to-Fine Retrieval for Camera Re-LocalizationMingyu Ding, Zhe Wang, Jiankai Sun, Jianping Shi 等ICCV 2019 · 被引用 163 次
- Expert Sample Consensus Applied to Camera Re-LocalizationEric Brachmann, Carsten RotherICCV 2019 · 被引用 136 次
- SANet: Scene Agnostic Network for Camera LocalizationLuwei Yang, Ziqian Bai, Chengzhou Tang, Honghua Li 等ICCV 2019 · 被引用 105 次
- Fine-Grained Segmentation Networks: Self-Supervised Segmentation for Improved Long-Term Visual LocalizationMåns Larsson, Erik Stenborg, Carl Toft, Lars Hammarstrand 等ICCV 2019 · 被引用 75 次
相关 Paper
- Learning Camera Localization via Dense Scene MatchingShitao Tang, Chengzhou Tang, Rui Huang, Siyu Zhu 等CVPR 2021
- Back to the Feature: Learning Robust Camera Localization From Pixels To PosePaul-Edouard Sarlin, Ajaykumar Unagar, Måns Larsson, Hugo Germain 等CVPR 2021
- VS-Net: Voting With Segmentation for Visual LocalizationZhaoyang Huang, Han Zhou, Yijin Li, Bangbang Yang 等CVPR 2021
- UprightNet: Geometry-Aware Camera Orientation Estimation From Single ImagesWenqi Xian, Zhengqi Li, Noah Snavely, Matthew Fisher 等ICCV 2019 · 被引用 52 次
- GLACE: Global Local Accelerated Coordinate EncodingFangjinhua Wang, Xudong Jiang, Silvano Galliani, Christoph Vogel 等CVPR 2024 · 被引用 24 次
