StructVPR: Distill Structural Knowledge with Weighting Samples for Visual Place Recognition
Yanqing Shen, Sanping Zhou, Jingwen Fu, Ruotong Wang, Shitao Chen, Nanning Zheng
Abstract
Visual place recognition (VPR) is usually considered as a specific image retrieval problem. Limited by existing training frameworks, most deep learning-based works cannot extract sufficiently stable global features from RGB images and rely on a time-consuming re-ranking step to exploit spatial structural information for better performance. In this paper, we propose StructVPR, a novel training architecture for VPR, to enhance structural knowledge in RGB global features and thus improve feature stability in a constantly changing environment. Specifically, StructVPR uses segmentation images as a more definitive source of structural knowledge input into a CNN network and applies knowledge distillation to avoid online segmentation and inference of seg-branch in testing. Considering that not all samples contain high-quality and helpful knowledge, and some even hurt the performance of distillation, we partition samples and weigh each sample's distillation loss to enhance the expected knowledge precisely. Finally, StructVPR achieves impressive performance on several benchmarks using only global retrieval and even outperforms many twostage approaches by a large margin. After adding additional re-ranking, ours achieves state-of-the-art performance while maintaining a low computational cost.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bd794b67-5749-47c3-9739-6a975d76ecebCited by top-tier papers5
- Towards Seamless Adaptation of Pre-trained Models for Visual Place RecognitionFeng Lu, Lijun Zhang, Xiangyuan Lan, Shuting Dong et al.ICLR 2024 · 81 citations
- Towards Implicit Aggregation: Robust Image Representation for Place Recognition in the Transformer EraFeng Lu, Tong Jin, Canming Ye, Xiangyuan Lan et al.NeurIPS 2025 · 8 citations
- Efficient Visual Place Recognition Through Multimodal Semantic Knowledge IntegrationSitao Zhang, Hongda Mao, Qingshuang Chen, Yelin KimICCV 2025 · 4 citations
- TopicGeo: An Efficient Unified Framework for GeolocationXin Wang, Xinlin Wang, Shuiping GouICCV 2025 · 2 citations
- HypeVPR: Exploring Hyperbolic Space for Perspective to Equirectangular Visual Place RecognitionSuhan Woo, Seongwon Lee, jinwoo jang, Euntai KimCVPR 2026
Builds on11
- Learning With Average Precision: Training Image Retrieval With a Listwise LossJérôme Revaud, Jon Almazán, Rafael S. Rezende, César Roberto de SouzaICCV 2019 · 424 citations
- TransVPR: Transformer-Based Place Recognition with Multi-Level Attention AggregationRuotong Wang, Yanqing Shen, Weiliang Zuo, Sanping Zhou et al.CVPR 2022 · 167 citations
- Discriminative Feature Learning With Consistent Attention Regularization for Person Re-IdentificationSanping Zhou, Fei Wang, Zeyi Huang, Jinjun WangICCV 2019 · 106 citations
- Learning an Augmented RGB Representation with Cross-Modal Knowledge Distillation for Action DetectionRui Dai, Srijan Das, François BrémondICCV 2021 · 50 citations
- SuperGlue: Learning Feature Matching With Graph Neural NetworksPaul-Edouard Sarlin, Daniel DeTone, Tomasz Malisiewicz, Andrew RabinovichCVPR 2020
Related papers
- D²-VPR: A Parameter-efficient Visual-foundation-model-based Visual Place Recognition Method via Knowledge Distillation and Deformable AggregationZheyuan Zhang, Jiwei Zhang, Boyu Zhou, Linzhimeng Duan et al.AAAI 2026 · 2 citations
- CricaVPR: Cross-Image Correlation-Aware Representation Learning for Visual Place RecognitionFeng Lu, Xiangyuan Lan, Lijun Zhang, Dongmei Jiang et al.CVPR 2024 · 68 citations
- DistilVPR: Cross-Modal Knowledge Distillation for Visual Place RecognitionSijie Wang, Rui She, Qiyu Kang, Xingchao Jian et al.AAAI 2024 · 14 citations
- Focus on Local: Finding Reliable Discriminative Regions for Visual Place RecognitionChangwei Wang, Shunpeng Chen, Yukun Song, Rongtao Xu et al.AAAI 2025 · 24 citations
- Deep Homography Estimation for Visual Place RecognitionFeng Lu, Shuting Dong, Lijun Zhang, Bingxi Liu et al.AAAI 2024 · 23 citations
