Patch-NetVLAD: Multi-Scale Fusion of Locally-Global Descriptors for Place Recognition
Stephen Hausler, Sourav Garg, Ming Xu, Michael Milford, Tobias Fischer
Abstract
Visual Place Recognition is a challenging task for robotics and autonomous systems, which must deal with the twin problems of appearance and viewpoint change in an always changing world. This paper introduces Patch-NetVLAD, which provides a novel formulation for combining the advantages of both local and global descriptor methods by deriving patch-level features from NetVLAD residuals. Unlike the fixed spatial neighborhood regime of existing local keypoint features, our method enables aggregation and matching of deep-learned local features defined over the feature-space grid. We further introduce a multi-scale fusion of patch features that have complementary scales (i.e. patch sizes) via an integral feature space and show that the fused features are highly invariant to both condition (season, structure, and illumination) and viewpoint (translation and rotation) changes. Patch-NetVLAD outperforms both global and local feature descriptor-based methods with comparable compute, achieving state-of-the-art visual place recognition results on a range of challenging real-world datasets, including winning the Facebook Mapillary Visual Place Recognition Challenge at ECCV2020. It is also adaptable to user requirements, with a speed-optimised version operating over an order of magnitude faster than the stateof-the-art. By combining superior performance with improved computational efficiency in a configurable framework, Patch-NetVLAD is well suited to enhance both stand-alone place recognition capabilities and the overall performance of SLAM systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c06a1f29-c6fd-44a4-916f-9854c42fc6b0Cited by top-tier papers59
- Rethinking Visual Geo-localization for Large-Scale ApplicationsGabriele Moreno Berton, Carlo Masone, Barbara CaputoCVPR 2022 · 235 citations
- TransVPR: Transformer-Based Place Recognition with Multi-Level Attention AggregationRuotong Wang, Yanqing Shen, Weiliang Zuo, Sanping Zhou et al.CVPR 2022 · 167 citations
- EigenPlaces: Training Viewpoint Robust Models for Visual Place RecognitionGabriele Moreno Berton, Gabriele Trivigno, Barbara Caputo, Carlo MasoneICCV 2023 · 141 citations
- Efficient LoFTR: Semi-Dense Local Feature Matching with Sparse-Like SpeedYifan Wang, Xingyi He, Sida Peng, Dongli Tan et al.CVPR 2024 · 126 citations
- SVT-Net: Super Light-Weight Sparse Voxel Transformer for Large Scale Place RecognitionZhaoxin Fan, Zhenbo Song, Hongyan Liu, Zhiwu Lu et al.AAAI 2022 · 95 citations
Builds on10
- Learning With Average Precision: Training Image Retrieval With a Listwise LossJérôme Revaud, Jon Almazán, Rafael S. Rezende, César Roberto de SouzaICCV 2019 · 424 citations
- Learning Two-View Correspondences and Geometry Using Order-Aware NetworkJiahui Zhang, Dawei Sun, Zixin Luo, Anbang Yao et al.ICCV 2019 · 362 citations
- LPD-Net: 3D Point Cloud Learning for Large-Scale Place Recognition and Environment AnalysisZhe Liu, Shunbo Zhou, Chuanzhe Suo, Peng Yin et al.ICCV 2019 · 337 citations
- Scalable Place Recognition Under Appearance Change for Autonomous DrivingDzung A. Doan, Yasir Latif, Tat-Jun Chin, Yu Liu et al.ICCV 2019 · 81 citations
- TextPlace: Visual Place Recognition and Topological Localization Through Reading Scene TextsZiyang Hong, Yvan R. Petillot, David Lane, Yishu Miao et al.ICCV 2019 · 60 citations
Related papers
- BEVPlace: Learning LiDAR-based Place Recognition using Bird's Eye View ImagesLun Luo, Shuhang Zheng, Yixuan Li, Yongzhi Fan et al.ICCV 2023 · 97 citations
- SuperVLAD: Compact and Robust Image Descriptors for Visual Place RecognitionFeng Lu, Xinyao Zhang, Canming Ye, Shuting Dong et al.NeurIPS 2024 · 24 citations
- Towards Implicit Aggregation: Robust Image Representation for Place Recognition in the Transformer EraFeng Lu, Tong Jin, Canming Ye, Xiangyuan Lan et al.NeurIPS 2025 · 8 citations
- Optimal Transport Aggregation for Visual Place RecognitionSergio Izquierdo, Javier CiveraCVPR 2024
- CricaVPR: Cross-Image Correlation-Aware Representation Learning for Visual Place RecognitionFeng Lu, Xiangyuan Lan, Lijun Zhang, Dongmei Jiang et al.CVPR 2024 · 68 citations
