VS-Net: Voting With Segmentation for Visual Localization
Zhaoyang Huang, Han Zhou, Yijin Li, Bangbang Yang, Yan Xu, Xiaowei Zhou, Hujun Bao, Guofeng Zhang, Hongsheng Li
摘要
Visual localization is of great importance in robotics and computer vision. Recently, scene coordinate regression based methods have shown good performance in visual localization in small static scenes. However, it still estimates camera poses from many inferior scene coordinates. To address this problem, we propose a novel visual localization framework that establishes 2D-to-3D correspondences between the query image and the 3D map with a series of learnable scene-specific landmarks. In the landmark generation stage, the 3D surfaces of the target scene are oversegmented into mosaic patches whose centers are regarded as the scene-specific landmarks. To robustly and accurately recover the scene-specific landmarks, we propose the Voting with Segmentation Network (VS-Net) to segment the pixels into different landmark patches with a segmentation branch and estimate the landmark locations within each patch with a landmark location voting branch. Since the number of landmarks in a scene may reach up to 5000, training a segmentation network with such a large number of classes is both computation and memory costly for the commonly used cross-entropy loss. We propose a novel prototype-based triplet loss with hard negative mining, which is able to train semantic segmentation networks with a large number of labels efficiently. Our proposed VS-Net is extensively tested on multiple public benchmarks and can outperform stateof-the-art visual localization methods. Code and models are available at https://github.com/zju3dv/VS-Net .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- RNNPose: Recurrent 6-DoF Object Pose Refinement with Robust Correspondence Field Estimation and Pose OptimizationYan Xu, Kwan-Yee Lin, Guofeng Zhang, Xiaogang Wang 等CVPR 2022 · 被引用 82 次
- Efficient Large-scale Localization by Global Instance RecognitionFei Xue, Ignas Budvytis, Daniel Olmeda Reino, Roberto CipollaCVPR 2022 · 被引用 19 次
- OFVL-MS: Once for Visual Localization across Multiple Indoor ScenesTao Xie, Kun Dai, Siyi Lu, Ke Wang 等ICCV 2023 · 被引用 16 次
- ETO: Efficient Transformer-based Local Feature Matching by Organizing Multiple Homography HypothesesJunjie Ni, Guofeng Zhang, Guanglin Li, Yijin Li 等NeurIPS 2024 · 被引用 14 次
- A Global Depth-Range-Free Multi-View Stereo Transformer Network with Pose EmbeddingYitong Dong, Yijin Li, Zhaoyang Huang, Weikang Bian 等NeurIPS 2024 · 被引用 7 次
它引用的顶会 Paper11
- SoftTriple Loss: Deep Metric Learning Without Triplet SamplingQi Qian, Lei Shang, Baigui Sun, Juhua Hu 等ICCV 2019 · 被引用 419 次
- AtLoc: Attention Guided Camera LocalizationBing Wang, Changhao Chen, Chris Xiaoxuan Lu, Peijun Zhao 等AAAI 2020 · 被引用 189 次
- Expert Sample Consensus Applied to Camera Re-LocalizationEric Brachmann, Carsten RotherICCV 2019 · 被引用 136 次
- Prior Guided Dropout for Robust Visual Localization in Dynamic EnvironmentsZhaoyang Huang, Yan Xu, Jianping Shi, Xiaowei Zhou 等ICCV 2019 · 被引用 54 次
- Squeeze-and-Attention Networks for Semantic SegmentationZilong Zhong, Zhong Qiu Lin, Rene Bidart, Xiaodan Hu 等CVPR 2020
相关 Paper
- Hierarchical Scene Coordinate Classification and Regression for Visual LocalizationXiaotian Li, Shuzhe Wang, Yi Zhao, Jakob Verbeek 等CVPR 2020
- SegLoc: Learning Segmentation-Based Representations for Privacy-Preserving Visual LocalizationMaxime Pietrantoni, Martin Humenberger, Torsten Sattler, Gabriela CsurkaCVPR 2023
- DenserNet: Weakly Supervised Visual Localization Using Multi-Scale Feature AggregationDongfang Liu, Yiming Cui, Liqi Yan, Christos Mousas 等AAAI 2021 · 被引用 149 次
- Reloc3r: Large-Scale Training of Relative Camera Pose Regression for Generalizable, Fast, and Accurate Visual LocalizationSiyan Dong, Shuzhe Wang, Shaohui Liu, Lulu Cai 等CVPR 2025
- EP2P-Loc: End-to-End 3D Point to 2D Pixel Localization for Large-Scale Visual LocalizationMinjung Kim, Junseo Koo, Gunhee KimICCV 2023 · 被引用 22 次
