EffoVPR: Effective Foundation Model Utilization for Visual Place Recognition
Issar Tzachor, Boaz Lerner, Matan Levy, Michael Green, Tal Berkovitz Shalev, Gavriel Habib, Dvir Samuel, Noam Korngut Zailer, Or Shimshi, Nir Darshan, Rami Ben-Ari
Abstract
The task of Visual Place Recognition (VPR) is to predict the location of a query image from a database of geo-tagged images. Recent studies in VPR have highlighted the significant advantage of employing pre-trained foundation models like DINOv2 for the VPR task. However, these models are often deemed inadequate for VPR without further fine-tuning on VPR-specific data. In this paper, we present an effective approach to harness the potential of a foundation model for VPR. We show that features extracted from self-attention layers can act as a powerful re-ranker for VPR, even in a zero-shot setting. Our method not only outperforms previous zero-shot approaches but also introduces results competitive with several supervised methods. We then show that a single-stage approach utilizing internal ViT layers for pooling can produce global features that achieve state-of-the-art performance, with impressive feature compactness down to 128D. Moreover, integrating our local foundation features for re-ranking further widens this performance gap. Our method also demonstrates exceptional robustness and generalization, setting new state-of-the-art performance, while handling challenging conditions such as occlusion, day-night transitions, and seasonal variations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8ea076ae-06e4-4985-b3ae-98d211512520Cited by top-tier papers6
- SAGE: Spatial-visual Adaptive Graph Exploration for Efficient Visual Place RecognitionShunpeng Chen, Changwei Wang, Rongtao Xu, Xingtian Pei et al.ICLR 2026 · 6 citations
- Vpr-Cloak: a First Look at Privacy Cloak Against Visual Place RecognitionShuting Dong, Mingzhi Chen, Feng Lu, Hao Yu et al.ICCV 2025 · 2 citations
- D²-VPR: A Parameter-efficient Visual-foundation-model-based Visual Place Recognition Method via Knowledge Distillation and Deformable AggregationZheyuan Zhang, Jiwei Zhang, Boyu Zhou, Linzhimeng Duan et al.AAAI 2026 · 2 citations
- Find your Needle: Small Object Image Retrieval via Multi-Object Attention OptimizationMichael Green, Matan Levy, Issar Tzachor, Dvir Samuel et al.NeurIPS 2025 · 1 citation
- CarGait: Cross-Attention Based Re-ranking for Gait RecognitionGavriel Habib, Noa Barzilay, Or Shimshi, Rami Ben-Ari et al.ICCV 2025
Builds on24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
Related papers
- EfficientVPR: Toward Efficient Visual Place Recognition via Scene-Aware Prompt Tuning and Adaptive Feature EnhancementWenjing Tang, Chuanguang Yang, Zhulin An, Libo Huang et al.CVPR 2026
- MutualVPR: A Mutual Learning Framework for Resolving Supervision Inconsistencies via Adaptive ClusteringQiwen Gu, Xufei Wang, Junqiao Zhao, Siyue Tao et al.NeurIPS 2025 · 6 citations
- TransVPR: Transformer-Based Place Recognition with Multi-Level Attention AggregationRuotong Wang, Yanqing Shen, Weiliang Zuo, Sanping Zhou et al.CVPR 2022 · 167 citations
- Towards Test-time Efficient Visual Place Recognition via Asymmetric Query ProcessingJaeyoon Kim, Yoonki Cho, Sung-Eui YoonAAAI 2026
- Optimal Transport Aggregation for Visual Place RecognitionSergio Izquierdo, Javier CiveraCVPR 2024
