CricaVPR: Cross-Image Correlation-Aware Representation Learning for Visual Place Recognition
Feng Lu, Xiangyuan Lan, Lijun Zhang, Dongmei Jiang, Yaowei Wang, Chun Yuan
Abstract
Over the past decade, most methods in visual place recognition (VPR) have used neural networks to produce feature representations. These networks typically produce a global representation of a place image using only this image itself and neglect the cross-image variations (e.g. viewpoint and illumination), which limits their robustness in challenging scenes. In this paper, we propose a robust global representation method with cross-image correlation awareness for VPR, named CricaVPR. Our method uses the attention mechanism to correlate multiple images within a batch. These images can be taken in the same place with different conditions or viewpoints, or even captured from different places. Therefore, our method can utilize the cross-image variations as a cue to guide the representation learning, which ensures more robust features are produced. To further facilitate the robustness, we propose a multi-scale convolution-enhanced adaptation method to adapt pre-trained visual foundation models to the VPR task, which introduces the multi-scale local information to further enhance the cross-image correlation-aware representation. Experimental results show that our method out-performs state-of-the-art methods by a large margin with significantly less training time. The code is released at https://github.com/Lu-Feng/CricaVPR.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7805dd0a-d7c5-4801-a48f-b5480bcf5213Cited by top-tier papers21
- Recognition through Reasoning: Reinforcing Image Geo-localization with Large Vision-Language ModelsLing Li, Yao Zhou, Yuxuan Liang, Fugee Tsung et al.NeurIPS 2025 · 30 citations
- SuperVLAD: Compact and Robust Image Descriptors for Visual Place RecognitionFeng Lu, Xinyao Zhang, Canming Ye, Shuting Dong et al.NeurIPS 2024 · 24 citations
- Focus on Local: Finding Reliable Discriminative Regions for Visual Place RecognitionChangwei Wang, Shunpeng Chen, Yukun Song, Rongtao Xu et al.AAAI 2025 · 24 citations
- GIM: A Million-scale Benchmark for Generative Image Manipulation Detection and LocalizationYirui Chen, Xudong Huang, Quan Zhang, Wei Li et al.AAAI 2025 · 15 citations
- Rethinking Pseudo-Label Guided Learning for Weakly Supervised Temporal Action Localization from the Perspective of Noise CorrectionQuan Zhang, Yuxin Qi, Xi Tang, Rui Yuan et al.AAAI 2025 · 11 citations
Builds on23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- AdaptFormer: Adapting Vision Transformers for Scalable Visual RecognitionShoufa Chen, Chongjian Ge, Zhan Tong, Jiangliu Wang et al.NeurIPS 2022 · 1,291 citations
- ST-Adapter: Parameter-Efficient Image-to-Video Transfer LearningJunting Pan, Ziyi Lin, Xiatian Zhu, Jing Shao et al.NeurIPS 2022 · 290 citations
Related papers
- Towards Seamless Adaptation of Pre-trained Models for Visual Place RecognitionFeng Lu, Lijun Zhang, Xiangyuan Lan, Shuting Dong et al.ICLR 2024 · 81 citations
- EigenPlaces: Training Viewpoint Robust Models for Visual Place RecognitionGabriele Moreno Berton, Gabriele Trivigno, Barbara Caputo, Carlo MasoneICCV 2023 · 141 citations
- TransVPR: Transformer-Based Place Recognition with Multi-Level Attention AggregationRuotong Wang, Yanqing Shen, Weiliang Zuo, Sanping Zhou et al.CVPR 2022 · 167 citations
- EffoVPR: Effective Foundation Model Utilization for Visual Place RecognitionIssar Tzachor, Boaz Lerner, Matan Levy, Michael Green et al.ICLR 2025
- StructVPR: Distill Structural Knowledge with Weighting Samples for Visual Place RecognitionYanqing Shen, Sanping Zhou, Jingwen Fu, Ruotong Wang et al.CVPR 2023
