EMVP: Embracing Visual Foundation Model for Visual Place Recognition with Centroid-Free Probing
Qibo Qiu, Shun Zhang, Haiming Gao, Honghui Yang, Haochao Ying, Wenxiao Wang, Xiaofei He
Abstract
Visual Place Recognition (VPR) is essential for mobile robots as it enables them to retrieve images from a database closest to their current location. The progress of Visual Foundation Models (VFMs) has significantly advanced VPR by capturing representative descriptors in images. However, existing fine-tuning efforts for VFMs often overlook the crucial role of probing in effectively adapting these descriptors for improved image representation. In this paper, we propose the Centroid-Free Probing (CFP) stage, making novel use of second-order features for more effective use of descriptors from VFMs. Moreover, to control the preservation of task-specific information adaptively based on the context of the VPR, we introduce the Dynamic Power Normalization (DPN) module in both the recalibration and CFP stages, forming a novel Parameter Efficiency Fine-Tuning (PEFT) pipeline (EMVP) tailored for the VPR task. Extensive experiments demonstrate the superiority of the proposed CFP over existing probing methods. Moreover, the EMVP pipeline can further enhance fine-tuning performance in terms of accuracy and efficiency. Specifically, it achieves 93.9%, 96.5%, and 94.6% Recall@1 on the MSLS Validation, Pitts250k-test, and SPED datasets, respectively, while saving 64.3% of trainable parameters compared with the existing SOTA PEFT method. The code is available at https://github.com/vincentqqb/EMVP .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- SAGE: Spatial-visual Adaptive Graph Exploration for Efficient Visual Place RecognitionShunpeng Chen, Changwei Wang, Rongtao Xu, Xingtian Pei et al.ICLR 2026 · 6 citations
- Vpr-Cloak: a First Look at Privacy Cloak Against Visual Place RecognitionShuting Dong, Mingzhi Chen, Feng Lu, Hao Yu et al.ICCV 2025 · 2 citations
- DialogueVPR: Towards Conversational Visual Place RecognitionYukun Song, Changwei Wang, Xingtian Pei, Shibiao Xu et al.CVPR 2026 · 1 citation
- Towards Test-time Efficient Visual Place Recognition via Asymmetric Query ProcessingJaeyoon Kim, Yoonki Cho, Sung-Eui YoonAAAI 2026
Builds on21
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
Related papers
- Tuning Pre-trained Model via Moment ProbingMingze Gao, Qilong Wang, Zhenyi Lin, Pengfei Zhu et al.ICCV 2023 · 16 citations
- Dynamic Tuning Towards Parameter and Inference Efficiency for ViT AdaptationWangbo Zhao, Jiasheng Tang, Yizeng Han, Yibing Song et al.NeurIPS 2024 · 41 citations
- EfficientVPR: Toward Efficient Visual Place Recognition via Scene-Aware Prompt Tuning and Adaptive Feature EnhancementWenjing Tang, Chuanguang Yang, Zhulin An, Libo Huang et al.CVPR 2026
- EffoVPR: Effective Foundation Model Utilization for Visual Place RecognitionIssar Tzachor, Boaz Lerner, Matan Levy, Michael Green et al.ICLR 2025
- CVPT: Cross Visual Prompt TuningLingyun Huang, Jianxu Mao, Junfei Yi, Ziming Tao et al.ICCV 2025 · 6 citations
