Rethinking Visual Geo-localization for Large-Scale Applications
Gabriele Moreno Berton, Carlo Masone, Barbara Caputo
摘要
Visual Geo-localization (VG) is the task of estimating the position where a given photo was taken by comparing it with a large database of images of known locations. To investigate how existing techniques would perform on a real-world city-wide VG application, we build San Francisco eXtra Large, a new dataset covering a whole city and providing a wide range of challenging cases, with a size 30x bigger than the previous largest dataset for visual geo-localization. We find that current methods fail to scale to such large datasets, therefore we design a new highly scalable training technique, called CosPlace, which casts the training as a classification problem avoiding the expensive mining needed by the commonly used contrastive learning. We achieve state-of-the-art performance on a wide range of datasets and find that CosPlace is robust to heavy domain changes. Moreover, we show that, compared to the previous state-of-the-art, CosPlace requires roughly 80% less GPU memory at train time, and it achieves better results with 8x smaller descriptors, paving the way for city-wide real-world visual geo-localization. Dataset, code and trained models are available for research purposes at https://github.com/gmberton/CosPlace .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper55
- EigenPlaces: Training Viewpoint Robust Models for Visual Place RecognitionGabriele Moreno Berton, Gabriele Trivigno, Barbara Caputo, Carlo MasoneICCV 2023 · 被引用 141 次
- Remote Sensing Vision-Language Foundation Models without Annotations via Ground Remote AlignmentUtkarsh Mall, Cheng Perng Phoo, Meilin Kelsey Liu, Carl Vondrick 等ICLR 2024 · 被引用 90 次
- Towards Seamless Adaptation of Pre-trained Models for Visual Place RecognitionFeng Lu, Lijun Zhang, Xiangyuan Lan, Shuting Dong 等ICLR 2024 · 被引用 81 次
- CricaVPR: Cross-Image Correlation-Aware Representation Learning for Visual Place RecognitionFeng Lu, Xiangyuan Lan, Lijun Zhang, Dongmei Jiang 等CVPR 2024 · 被引用 68 次
- Recognition through Reasoning: Reinforcing Image Geo-localization with Large Vision-Language ModelsLing Li, Yao Zhou, Yuxuan Liang, Fugee Tsung 等NeurIPS 2025 · 被引用 30 次
它引用的顶会 Paper10
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Stochastic Attraction-Repulsion Embedding for Large Scale Image LocalizationLiu Liu, Hongdong Li, Yuchao DaiICCV 2019 · 被引用 123 次
- Instance-level Image Retrieval using Reranking TransformersFuwen Tan, Jiangbo Yuan, Vicente OrdonezICCV 2021 · 被引用 116 次
- Scalable Place Recognition Under Appearance Change for Autonomous DrivingDzung A. Doan, Yasir Latif, Tat-Jun Chin, Yu Liu 等ICCV 2019 · 被引用 81 次
- Deep Visual Geo-localization BenchmarkGabriele Moreno Berton, Riccardo Mereu, Gabriele Trivigno, Carlo Masone 等CVPR 2022 · 被引用 80 次
相关 Paper
- HierLoc: Hyperbolic Entity Embeddings for Hierarchical Visual GeolocationHari Krishna Gadi, Daniel Matos, Hongyi Luo, Lu Liu 等ICLR 2026 · 被引用 1 次
- Data-Efficient Large Scale Place Recognition with Graded Similarity SupervisionMaria Leyva-Vallina, Nicola Strisciuglio, Nicolai PetkovCVPR 2023
- CrossLoc: Scalable Aerial Localization Assisted by Multimodal Synthetic DataQi Yan, Jianhao Zheng, Simon Reding, Shanci Li 等CVPR 2022 · 被引用 23 次
- Soft Contrastive Learning for Visual LocalizationJanine Thoma, Danda Pani Paudel, Luc Van GoolNeurIPS 2020 · 被引用 41 次
- Sample4Geo: Hard Negative Sampling For Cross-View Geo-LocalisationFabian Deuser, Konrad Habel, Norbert OswaldICCV 2023 · 被引用 161 次
