VS-Net: Voting With Segmentation for Visual Localization
Zhaoyang Huang, Han Zhou, Yijin Li, Bangbang Yang, Yan Xu, Xiaowei Zhou, Hujun Bao, Guofeng Zhang, Hongsheng Li
Abstract
Visual localization is of great importance in robotics and computer vision. Recently, scene coordinate regression based methods have shown good performance in visual localization in small static scenes. However, it still estimates camera poses from many inferior scene coordinates. To address this problem, we propose a novel visual localization framework that establishes 2D-to-3D correspondences between the query image and the 3D map with a series of learnable scene-specific landmarks. In the landmark generation stage, the 3D surfaces of the target scene are oversegmented into mosaic patches whose centers are regarded as the scene-specific landmarks. To robustly and accurately recover the scene-specific landmarks, we propose the Voting with Segmentation Network (VS-Net) to segment the pixels into different landmark patches with a segmentation branch and estimate the landmark locations within each patch with a landmark location voting branch. Since the number of landmarks in a scene may reach up to 5000, training a segmentation network with such a large number of classes is both computation and memory costly for the commonly used cross-entropy loss. We propose a novel prototype-based triplet loss with hard negative mining, which is able to train semantic segmentation networks with a large number of labels efficiently. Our proposed VS-Net is extensively tested on multiple public benchmarks and can outperform stateof-the-art visual localization methods. Code and models are available at https://github.com/zju3dv/VS-Net .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3c1d754f-dbef-41ce-a831-ae765d037d0cCited by top-tier papers10
- RNNPose: Recurrent 6-DoF Object Pose Refinement with Robust Correspondence Field Estimation and Pose OptimizationYan Xu, Kwan-Yee Lin, Guofeng Zhang, Xiaogang Wang et al.CVPR 2022 · 82 citations
- Efficient Large-scale Localization by Global Instance RecognitionFei Xue, Ignas Budvytis, Daniel Olmeda Reino, Roberto CipollaCVPR 2022 · 19 citations
- OFVL-MS: Once for Visual Localization across Multiple Indoor ScenesTao Xie, Kun Dai, Siyi Lu, Ke Wang et al.ICCV 2023 · 16 citations
- ETO: Efficient Transformer-based Local Feature Matching by Organizing Multiple Homography HypothesesJunjie Ni, Guofeng Zhang, Guanglin Li, Yijin Li et al.NeurIPS 2024 · 14 citations
- A Global Depth-Range-Free Multi-View Stereo Transformer Network with Pose EmbeddingYitong Dong, Yijin Li, Zhaoyang Huang, Weikang Bian et al.NeurIPS 2024 · 7 citations
Builds on11
- SoftTriple Loss: Deep Metric Learning Without Triplet SamplingQi Qian, Lei Shang, Baigui Sun, Juhua Hu et al.ICCV 2019 · 419 citations
- AtLoc: Attention Guided Camera LocalizationBing Wang, Changhao Chen, Chris Xiaoxuan Lu, Peijun Zhao et al.AAAI 2020 · 189 citations
- Expert Sample Consensus Applied to Camera Re-LocalizationEric Brachmann, Carsten RotherICCV 2019 · 136 citations
- Prior Guided Dropout for Robust Visual Localization in Dynamic EnvironmentsZhaoyang Huang, Yan Xu, Jianping Shi, Xiaowei Zhou et al.ICCV 2019 · 54 citations
- Squeeze-and-Attention Networks for Semantic SegmentationZilong Zhong, Zhong Qiu Lin, Rene Bidart, Xiaodan Hu et al.CVPR 2020
Related papers
- Hierarchical Scene Coordinate Classification and Regression for Visual LocalizationXiaotian Li, Shuzhe Wang, Yi Zhao, Jakob Verbeek et al.CVPR 2020
- SegLoc: Learning Segmentation-Based Representations for Privacy-Preserving Visual LocalizationMaxime Pietrantoni, Martin Humenberger, Torsten Sattler, Gabriela CsurkaCVPR 2023
- DenserNet: Weakly Supervised Visual Localization Using Multi-Scale Feature AggregationDongfang Liu, Yiming Cui, Liqi Yan, Christos Mousas et al.AAAI 2021 · 149 citations
- Reloc3r: Large-Scale Training of Relative Camera Pose Regression for Generalizable, Fast, and Accurate Visual LocalizationSiyan Dong, Shuzhe Wang, Shaohui Liu, Lulu Cai et al.CVPR 2025
- EP2P-Loc: End-to-End 3D Point to 2D Pixel Localization for Large-Scale Visual LocalizationMinjung Kim, Junseo Koo, Gunhee KimICCV 2023 · 22 citations
