Attentional Pyramid Pooling of Salient Visual Residuals for Place Recognition
Guohao Peng, Jun Zhang, Heshan Li, Danwei Wang
Abstract
The core of visual place recognition (VPR) lies in how to identify task-relevant visual cues and embed them into dis- criminative representations. Focusing on these two points, we propose a novel encoding strategy named Attentional Pyramid Pooling of Salient Visual Residuals (APPSVR). It incorporates three types of attention modules to model the saliency of local features in individual, spatial and cluster dimensions respectively. (1) To inhibit task-irrelevant local features, a semantic-reinforced local weighting scheme is employed for local feature refinement; (2) To leverage the spatial context, an attentional pyramid structure is constructed to adaptively encode regional features according to their relative spatial saliency; (3) To distinguish the different importance of visual clusters to the task, a parametric normalization is proposed to adjust their contribution to image descriptor generation. Experiments demonstrate APPSVR outperforms the existing techniques and achieves a new state-of-the-art performance on VPR benchmark datasets. The visualization shows the saliency map learned in a weakly supervised manner is largely consistent with human cognition.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers12
- Rethinking Visual Geo-localization for Large-Scale ApplicationsGabriele Moreno Berton, Carlo Masone, Barbara CaputoCVPR 2022 · 235 citations
- EigenPlaces: Training Viewpoint Robust Models for Visual Place RecognitionGabriele Moreno Berton, Gabriele Trivigno, Barbara Caputo, Carlo MasoneICCV 2023 · 141 citations
- Towards Seamless Adaptation of Pre-trained Models for Visual Place RecognitionFeng Lu, Lijun Zhang, Xiangyuan Lan, Shuting Dong et al.ICLR 2024 · 81 citations
- Deep Visual Geo-localization BenchmarkGabriele Moreno Berton, Riccardo Mereu, Gabriele Trivigno, Carlo Masone et al.CVPR 2022 · 80 citations
- CricaVPR: Cross-Image Correlation-Aware Representation Learning for Visual Place RecognitionFeng Lu, Xiangyuan Lan, Lijun Zhang, Dongmei Jiang et al.CVPR 2024 · 68 citations
Related papers
- Focus on Local: Finding Reliable Discriminative Regions for Visual Place RecognitionChangwei Wang, Shunpeng Chen, Yukun Song, Rongtao Xu et al.AAAI 2025 · 24 citations
- TransVPR: Transformer-Based Place Recognition with Multi-Level Attention AggregationRuotong Wang, Yanqing Shen, Weiliang Zuo, Sanping Zhou et al.CVPR 2022 · 167 citations
- Efficient Visual Place Recognition Through Multimodal Semantic Knowledge IntegrationSitao Zhang, Hongda Mao, Qingshuang Chen, Yelin KimICCV 2025 · 4 citations
- Towards Implicit Aggregation: Robust Image Representation for Place Recognition in the Transformer EraFeng Lu, Tong Jin, Canming Ye, Xiangyuan Lan et al.NeurIPS 2025 · 8 citations
- Pyramid Point Cloud Transformer for Large-Scale Place RecognitionLe Hui, Hang Yang, Mingmei Cheng, Jin Xie et al.ICCV 2021 · 147 citations
