Optimal Transport Aggregation for Visual Place Recognition
Sergio Izquierdo, Javier Civera
摘要
The task of Visual Place Recognition (VPR) aims to match a query image against references from an extensive database of images from different places, relying solely on visual cues. State-of-the-art pipelines focus on the aggregation of features extracted from a deep backbone, in order to form a global descriptor for each image. In this context, we introduce SALAD (Sinkhorn Algorithm for Locally Aggregated Descriptors), which reformulates NetVLAD's soft-assignment of local features to clusters as an optimal transport problem. In SALAD, we consider both featureto-cluster and cluster-to-feature relations and we also introduce a 'dustbin' cluster, designed to selectively discard features deemed non-informative, enhancing the overall descriptor quality. Additionally, we leverage and fine-tune DINOv2 as a backbone, which provides enhanced description power for the local features, and dramatically reduces the required training time. As a result, our single-stage method not only surpasses single-stage baselines in public VPR datasets, but also surpasses two-stage methods that add a re-ranking with significantly higher cost. Code and models are available at https://github.com/serizba/salad .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper39
- VGGT-SLAM: Dense RGB SLAM Optimized on the SL(4) ManifoldDominic Maggio, Hyungtae Lim, Luca CarloneNeurIPS 2025 · 被引用 176 次
- SuperVLAD: Compact and Robust Image Descriptors for Visual Place RecognitionFeng Lu, Xinyao Zhang, Canming Ye, Shuting Dong 等NeurIPS 2024 · 被引用 24 次
- Focus on Local: Finding Reliable Discriminative Regions for Visual Place RecognitionChangwei Wang, Shunpeng Chen, Yukun Song, Rongtao Xu 等AAAI 2025 · 被引用 24 次
- Multiview Scene GraphJuexiao Zhang, Gao Zhu, Sihang Li, Xinhao Liu 等NeurIPS 2024 · 被引用 13 次
- EMVP: Embracing Visual Foundation Model for Visual Place Recognition with Centroid-Free ProbingQibo Qiu, Shun Zhang, Haiming Gao, Honghui Yang 等NeurIPS 2024 · 被引用 13 次
它引用的顶会 Paper16
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer 等CVPR 2022 · 被引用 6,782 次
- Swin Transformer V2: Scaling Up Capacity and ResolutionZe Liu, Han Hu, Yutong Lin, Zhuliang Yao 等CVPR 2022 · 被引用 2,138 次
- MPViT: Multi-Path Vision Transformer for Dense PredictionYoungwan Lee, Jonghee Kim, Jeffrey Willette, Sung Ju HwangCVPR 2022 · 被引用 339 次
- Rethinking Visual Geo-localization for Large-Scale ApplicationsGabriele Moreno Berton, Carlo Masone, Barbara CaputoCVPR 2022 · 被引用 235 次
相关 Paper
- A2GC: Asymmetric Aggregation with Geometric Constraints for Locally Aggregated DescriptorsZhenyu Li, Tianyi ShangCVPR 2026 · 被引用 5 次
- Towards Implicit Aggregation: Robust Image Representation for Place Recognition in the Transformer EraFeng Lu, Tong Jin, Canming Ye, Xiangyuan Lan 等NeurIPS 2025 · 被引用 8 次
- EffoVPR: Effective Foundation Model Utilization for Visual Place RecognitionIssar Tzachor, Boaz Lerner, Matan Levy, Michael Green 等ICLR 2025
- SAGE: Spatial-visual Adaptive Graph Exploration for Efficient Visual Place RecognitionShunpeng Chen, Changwei Wang, Rongtao Xu, Xingtian Pei 等ICLR 2026 · 被引用 6 次
- EfficientVPR: Toward Efficient Visual Place Recognition via Scene-Aware Prompt Tuning and Adaptive Feature EnhancementWenjing Tang, Chuanguang Yang, Zhulin An, Libo Huang 等CVPR 2026
