CVGL: Causal Learning and Geometric Topology
Songsong Ouyang, Yingying Zhu
Abstract
Cross-view geo-localization (CVGL) aims to estimate the geographic location of a street image by matching it with a corresponding aerial image. This is critical for autonomous navigation and mapping in complex real-world scenarios. However, the task remains challenging due to significant viewpoint differences and the influence of confounding factors. To tackle these issues, we propose the Causal Learning and Geometric Topology (CLGT) framework, which integrates two key components: a Causal Feature Extractor (CFE) that mitigates the influence of confounding factors by leveraging causal intervention to encourage the model to focus on stable, task-relevant semantics; and a Geometric Topology Fusion (GT Fusion) module that injects Bird's Eye View (BEV) road topology into street features to alleviate cross-view inconsistencies caused by extreme perspective changes. Additionally, we introduce a Data-Adaptive Pooling (DA Pooling) module to enhance the representation of semantically rich regions. Extensive experiments on CVUSA, CVACT, and their robustness-enhanced variants (CVUSA-C-ALL and CVACT-C-ALL) demonstrate that CLGT achieves state-of-the-art performance, particularly under challenging real-world corruptions. Our codes are available at https://github.com/oyss-szu/CLGT.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 630b8aab-12e6-4e9f-ace5-ea3acb0cacb6Cited by top-tier papers2
- InfoGeo: Information-Theoretic Object-Centric Learning for Cross-View Generalizable UAV Geo-LocalizationHongyang ZHANG, Maonan Wang, Ziyao Wang, Hongrui Yin et al.ICML 2026 · 1 citation
- Must All Negatives Be Pushed Away Equally? Uncertainty-Aware Cross-View Geo-Localization via Normal Inverse Gamma DistributionSongsong Ouyang, Le Wu, Yingying ZhuICML 2026
Builds on18
- Long-Tailed Classification by Keeping the Good and Removing the Bad Momentum Causal EffectKaihua Tang, Jianqiang Huang, Hanwang ZhangNeurIPS 2020 · 533 citations
- Interventional Few-Shot LearningZhongqi Yue, Hanwang Zhang, Qianru Sun, Xian-Sheng HuaNeurIPS 2020 · 284 citations
- Cross-view Transformers for real-time Map-view Semantic SegmentationBrady Zhou, Philipp KrähenbühlCVPR 2022 · 279 citations
- Cross-view Geo-localization with Layer-to-Layer TransformerHongji Yang, Xiufan Lu, Yingying ZhuNeurIPS 2021 · 231 citations
- Optimal Feature Transport for Cross-View Image Geo-LocalizationYujiao Shi, Xin Yu, Liu Liu, Tong Zhang et al.AAAI 2020 · 210 citations
Related papers
- Geo2: Geometry-Guided Cross-view Geo-Localization and Image SynthesisYancheng Zhang, Xiaohan Zhang, Guangyu Sun, Zonglin Lyu et al.CVPR 2026 · 5 citations
- Cross-View Geo-Localization via Learning Disentangled Geometric Layout CorrespondenceXiaohan Zhang, Xingyu Li, Waqas Sultani, Yi Zhou et al.AAAI 2023 · 111 citations
- Uncertainty-Aware Vision-Based Metric Cross-View GeolocalizationFlorian Fervers, Sebastian Bullinger, Christoph Bodensteiner, Michael Arens et al.CVPR 2023
- Aligning Geometric Spatial Layout in Cross-View Geo-Localization via Feature RecombinationQingwang Zhang, Yingying ZhuAAAI 2024 · 28 citations
- Coming Down to Earth: Satellite-to-Street View Synthesis for Geo-LocalizationAysim Toker, Qunjie Zhou, Maxim Maximov, Laura Leal-TaixéCVPR 2021
