TransVPR: Transformer-Based Place Recognition with Multi-Level Attention Aggregation
Ruotong Wang, Yanqing Shen, Weiliang Zuo, Sanping Zhou, Nanning Zheng
摘要
Visual place recognition is a challenging task for applications such as autonomous driving navigation and mobile robot localization. Distracting elements presenting in complex scenes often lead to deviations in the perception of visual place. To address this problem, it is crucial to integrate information from only task-relevant regions into image representations. In this paper, we introduce a novel holistic place recognition model, TransVPR, based on vision Transformers. It benefits from the desirable property of the self-attention operation in Transformers which can naturally aggregate task-relevant features. Attentions from multiple levels of the Transformer, which focus on different regions of interest, are further combined to generate a global image representation. In addition, the output tokens from Transformer layers filtered by the fused attention mask are considered as key-patch descriptors, which are used to perform spatial matching to re-rank the candidates retrieved by the global image features. The whole model allows end-to-end training with a single objective and image-level supervision. TransVPR achieves state-of-the-art performance on several real-world benchmarks while maintaining low computational time and storage requirements.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper27
- EigenPlaces: Training Viewpoint Robust Models for Visual Place RecognitionGabriele Moreno Berton, Gabriele Trivigno, Barbara Caputo, Carlo MasoneICCV 2023 · 被引用 141 次
- Towards Seamless Adaptation of Pre-trained Models for Visual Place RecognitionFeng Lu, Lijun Zhang, Xiangyuan Lan, Shuting Dong 等ICLR 2024 · 被引用 81 次
- CricaVPR: Cross-Image Correlation-Aware Representation Learning for Visual Place RecognitionFeng Lu, Xiangyuan Lan, Lijun Zhang, Dongmei Jiang 等CVPR 2024 · 被引用 68 次
- SuperVLAD: Compact and Robust Image Descriptors for Visual Place RecognitionFeng Lu, Xinyao Zhang, Canming Ye, Shuting Dong 等NeurIPS 2024 · 被引用 24 次
- Focus on Local: Finding Reliable Discriminative Regions for Visual Place RecognitionChangwei Wang, Shunpeng Chen, Yukun Song, Rongtao Xu 等AAAI 2025 · 被引用 24 次
它引用的顶会 Paper5
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Learning With Average Precision: Training Image Retrieval With a Listwise LossJérôme Revaud, Jon Almazán, Rafael S. Rezende, César Roberto de SouzaICCV 2019 · 被引用 424 次
- SuperGlue: Learning Feature Matching With Graph Neural NetworksPaul-Edouard Sarlin, Daniel DeTone, Tomasz Malisiewicz, Andrew RabinovichCVPR 2020
- Mapillary Street-Level Sequences: A Dataset for Lifelong Place RecognitionFrederik Warburg, Søren Hauberg, Manuel López-Antequera, Pau Gargallo 等CVPR 2020
- Patch-NetVLAD: Multi-Scale Fusion of Locally-Global Descriptors for Place RecognitionStephen Hausler, Sourav Garg, Ming Xu, Michael Milford 等CVPR 2021
相关 Paper
- Towards Implicit Aggregation: Robust Image Representation for Place Recognition in the Transformer EraFeng Lu, Tong Jin, Canming Ye, Xiangyuan Lan 等NeurIPS 2025 · 被引用 8 次
- Former: Unified Retrieval and Reranking Transformer for Place RecognitionSijie Zhu, Linjie Yang, Chen Chen, Mubarak Shah 等CVPR 2023
- BoQ: A Place is Worth a Bag of Learnable QueriesAmar Ali-bey, Brahim Chaib-draa, Philippe GiguèreCVPR 2024
- Deep Homography Estimation for Visual Place RecognitionFeng Lu, Shuting Dong, Lijun Zhang, Bingxi Liu 等AAAI 2024 · 被引用 23 次
- EffoVPR: Effective Foundation Model Utilization for Visual Place RecognitionIssar Tzachor, Boaz Lerner, Matan Levy, Michael Green 等ICLR 2025
