SuperVLAD: Compact and Robust Image Descriptors for Visual Place Recognition
Feng Lu, Xinyao Zhang, Canming Ye, Shuting Dong, Lijun Zhang, Xiangyuan Lan, Chun Yuan
Abstract
Visual place recognition (VPR) is an essential task for multiple applications such as augmented reality and robot localization. Over the past decade, mainstream methods in the VPR area have been to use feature representation based on global aggregation, as exemplified by NetVLAD. These features are suitable for large-scale VPR and robust against viewpoint changes. However, the VLAD-based aggregation methods usually learn a large number of ( e.g. , 64) clusters and their corresponding cluster centers, which directly leads to a high dimension of the yielded global features. More importantly, when there is a domain gap between the data in training and inference, the cluster centers determined on the training set are usually improper for inference, resulting in a performance drop. To this end, we first attempt to improve NetVLAD by removing the cluster center and setting only a small number of ( e.g. , only 4) clusters. The proposed method not only simplifies NetVLAD but also enhances the generalizability across different domains. We name this method SuperVLAD . In addition, by introducing ghost clusters that will not be retained in the final output, we further propose a very low-dimensional 1-Cluster VLAD descriptor, which has the same dimension as the output of GeM pooling but performs notably better. Experimental results suggest that, when paired with a transformer-based backbone, our SuperVLAD shows better domain generalization performance than NetVLAD with significantly fewer parameters. The proposed method also surpasses state-of-the-art methods with lower feature dimensions on several benchmark datasets. The code is available at https://github.com/lu-feng/SuperVLAD.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 76184dd4-2b93-4c23-87cc-478ff55fdafdCited by top-tier papers6
- Towards Implicit Aggregation: Robust Image Representation for Place Recognition in the Transformer EraFeng Lu, Tong Jin, Canming Ye, Xiangyuan Lan et al.NeurIPS 2025 · 8 citations
- SAGE: Spatial-visual Adaptive Graph Exploration for Efficient Visual Place RecognitionShunpeng Chen, Changwei Wang, Rongtao Xu, Xingtian Pei et al.ICLR 2026 · 6 citations
- Vpr-Cloak: a First Look at Privacy Cloak Against Visual Place RecognitionShuting Dong, Mingzhi Chen, Feng Lu, Hao Yu et al.ICCV 2025 · 2 citations
- D²-VPR: A Parameter-efficient Visual-foundation-model-based Visual Place Recognition Method via Knowledge Distillation and Deformable AggregationZheyuan Zhang, Jiwei Zhang, Boyu Zhou, Linzhimeng Duan et al.AAAI 2026 · 2 citations
- DialogueVPR: Towards Conversational Visual Place RecognitionYukun Song, Changwei Wang, Xingtian Pei, Shibiao Xu et al.CVPR 2026 · 1 citation
Builds on20
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Rethinking Visual Geo-localization for Large-Scale ApplicationsGabriele Moreno Berton, Carlo Masone, Barbara CaputoCVPR 2022 · 235 citations
- TransVPR: Transformer-Based Place Recognition with Multi-Level Attention AggregationRuotong Wang, Yanqing Shen, Weiliang Zuo, Sanping Zhou et al.CVPR 2022 · 167 citations
- EigenPlaces: Training Viewpoint Robust Models for Visual Place RecognitionGabriele Moreno Berton, Gabriele Trivigno, Barbara Caputo, Carlo MasoneICCV 2023 · 141 citations
- Stochastic Attraction-Repulsion Embedding for Large Scale Image LocalizationLiu Liu, Hongdong Li, Yuchao DaiICCV 2019 · 123 citations
Related papers
- Optimal Transport Aggregation for Visual Place RecognitionSergio Izquierdo, Javier CiveraCVPR 2024
- Patch-NetVLAD: Multi-Scale Fusion of Locally-Global Descriptors for Place RecognitionStephen Hausler, Sourav Garg, Ming Xu, Michael Milford et al.CVPR 2021
- BEVPlace: Learning LiDAR-based Place Recognition using Bird's Eye View ImagesLun Luo, Shuhang Zheng, Yixuan Li, Yongzhi Fan et al.ICCV 2023 · 97 citations
- Pyramid Point Cloud Transformer for Large-Scale Place RecognitionLe Hui, Hang Yang, Mingmei Cheng, Jin Xie et al.ICCV 2021 · 147 citations
- MutualVPR: A Mutual Learning Framework for Resolving Supervision Inconsistencies via Adaptive ClusteringQiwen Gu, Xufei Wang, Junqiao Zhao, Siyue Tao et al.NeurIPS 2025 · 6 citations
