Cross-View Geo-Localization via Learning Disentangled Geometric Layout Correspondence
Xiaohan Zhang, Xingyu Li, Waqas Sultani, Yi Zhou, Safwan Wshah
摘要
Cross-view geo-localization aims to estimate the location of a query ground image by matching it to a reference geo-tagged aerial images database. As an extremely challenging task, its difficulties root in the drastic view changes and different capturing time between two views. Despite these difficulties, recent works achieve outstanding progress on cross-view geo-localization benchmarks. However, existing methods still suffer from poor performance on the cross-area benchmarks, in which the training and testing data are captured from two different regions. We attribute this deficiency to the lack of ability to extract the spatial configuration of visual feature layouts and models' overfitting on low-level details from the training set. In this paper, we propose GeoDTR which explicitly disentangles geometric information from raw features and learns the spatial correlations among visual features from aerial and ground pairs with a novel geometric layout extractor module. This module generates a set of geometric layout descriptors, modulating the raw features and producing high-quality latent representations. In addition, we elaborate on two categories of data augmentations, (i) Layout simulation, which varies the spatial configuration while keeping the low-level details intact. (ii) Semantic augmentation, which alters the low-level details and encourages the model to capture spatial configurations. These augmentations help to improve the performance of the cross-view geo-localization models, especially on the cross-area benchmarks. Moreover, we propose a counterfactual-based learning process to benefit the geometric layout extractor in exploring spatial information. Extensive experiments show that GeoDTR not only achieves state-of-the-art results but also significantly boosts the performance on same-area and cross-area benchmarks. Our code can be found at https://gitlab.com/vail-uvm/geodtr.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper23
- Sample4Geo: Hard Negative Sampling For Cross-View Geo-LocalisationFabian Deuser, Konrad Habel, Norbert OswaldICCV 2023 · 被引用 161 次
- GeoReasoner: Geo-localization with Reasoning in Street Views using a Large Vision-Language ModelLing Li, Yu Ye, Bingchuan Jiang, Wei ZengICML 2024 · 被引用 35 次
- Aligning Geometric Spatial Layout in Cross-View Geo-Localization via Feature RecombinationQingwang Zhang, Yingying ZhuAAAI 2024 · 被引用 28 次
- GeoRanker: Distance-Aware Ranking for Worldwide Image GeolocalizationPengyue Jia, Seongheon Park, Song Gao, Xiangyu Zhao 等NeurIPS 2025 · 被引用 22 次
- Why do Variational Autoencoders Really Promote Disentanglement?Pratik Bhowal, Achint Soni, Sirisha RambhatlaICML 2024 · 被引用 11 次
它引用的顶会 Paper9
- University-1652: A Multi-view Multi-source Benchmark for Drone-based Geo-localizationZhedong Zheng, Yunchao Wei, Yi YangACM MM 2020 · 被引用 390 次
- Counterfactual Attention Learning for Fine-Grained Visual Categorization and Re-identificationYongming Rao, Guangyi Chen, Jiwen Lu, Jie ZhouICCV 2021 · 被引用 330 次
- Cross-view Geo-localization with Layer-to-Layer TransformerHongji Yang, Xiufan Lu, Yingying ZhuNeurIPS 2021 · 被引用 231 次
- Bridging the Domain Gap for Ground-to-Aerial Image MatchingKrishna Regmi, Mubarak ShahICCV 2019 · 被引用 191 次
- Ground-to-Aerial Image Geo-Localization With a Hard Exemplar Reweighting Triplet LossSudong Cai, Yulan Guo, Salman H. Khan, Jiwei Hu 等ICCV 2019 · 被引用 140 次
相关 Paper
- VIGOR: Cross-View Image Geo-Localization Beyond One-to-One RetrievalSijie Zhu, Taojiannan Yang, Chen ChenCVPR 2021
- Geo2: Geometry-Guided Cross-view Geo-Localization and Image SynthesisYancheng Zhang, Xiaohan Zhang, Guangyu Sun, Zonglin Lyu 等CVPR 2026 · 被引用 5 次
- From Coarse to Fine: A Matching and Alignment Framework for Unsupervised Cross-View Geo-LocalizationXueyi Wang, Lele Zhang, Zheng Fan, Yang Liu 等AAAI 2025 · 被引用 12 次
- Game4Loc: A UAV Geo-Localization Benchmark from Game DataYuxiang Ji, Boyong He, Zhuoyue Tan, Liaoni WuAAAI 2025 · 被引用 35 次
- Cross-View Visual Geo-Localization for Outdoor Augmented RealityNiluthpol Chowdhury Mithun, Kshitij Minhas, Han-Pang Chiu, Taragay Oskiper 等IEEE VR 2023 · 被引用 22 次
