Efficient Visual Place Recognition Through Multimodal Semantic Knowledge Integration
Sitao Zhang, Hongda Mao, Qingshuang Chen, Yelin Kim
Abstract
Visual place recognition is crucial for autonomous navigation and robotic mapping. Current methods struggle with perceptual aliasing and computational inefficiency. We present SemVPR, a novel approach integrating multimodal semantic knowledge into VPR. By leveraging a pre-trained vision-language model as a teacher during the training phase, SemVPR learns local visual and semantic descriptors simultaneously, effectively mitigating perceptual aliasing through semantic-aware aggregation without extra inference cost. The proposed nested descriptor learning strategy generates a series of ultra-compact global descriptors, reduced by approximately 66× compared to state-of-the-art methods, in a coarse-to-fine manner, eliminating the need for offline dimensionality reduction or training multiple models. Extensive experiments across various VPR benchmarks demonstrate that SemVPR consistently outperforms state-of-the-art methods with significantly lower computational costs, rendering its feasibility for latency-sensitive scenarios in real-world applications.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2cf58e75-8da9-4710-86ed-9f297861e6efCited by top-tier papers1
Ask how each one uses itBuilds on20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
Related papers
- Towards Seamless Adaptation of Pre-trained Models for Visual Place RecognitionFeng Lu, Lijun Zhang, Xiangyuan Lan, Shuting Dong et al.ICLR 2024 · 81 citations
- TransVPR: Transformer-Based Place Recognition with Multi-Level Attention AggregationRuotong Wang, Yanqing Shen, Weiliang Zuo, Sanping Zhou et al.CVPR 2022 · 167 citations
- A Hyperdimensional One Place Signature to Represent Them All: Stackable Descriptors for Visual Place RecognitionConnor Malone, Somayeh Hussaini, Tobias Fischer, Michael MilfordICCV 2025 · 2 citations
- Towards Implicit Aggregation: Robust Image Representation for Place Recognition in the Transformer EraFeng Lu, Tong Jin, Canming Ye, Xiangyuan Lan et al.NeurIPS 2025 · 8 citations
- CricaVPR: Cross-Image Correlation-Aware Representation Learning for Visual Place RecognitionFeng Lu, Xiangyuan Lan, Lijun Zhang, Dongmei Jiang et al.CVPR 2024 · 68 citations
