Geo2: Geometry-Guided Cross-view Geo-Localization and Image Synthesis
Yancheng Zhang, Xiaohan Zhang, Guangyu Sun, Zonglin Lyu, Safwan Wshah, Chen Chen
摘要
Cross-view geo-spatial learning consists of two important tasks: Cross-View Geo-Localization (CVGL) and Cross-View Image Synthesis (CVIS), both of which rely on establishing geometric correspondences between ground and aerial views. Recent Geometric Foundation Models (GFMs) have demonstrated strong capabilities in extracting generalizable 3D geometric features from images, but their potential in cross-view geo-spatial tasks remains underexplored. In this work, we present Geo 2 , a unified framework that leverages Geometric priors from GFMs (e.g., VGGT) to jointly perform Geo-spatial tasks, CVGL and bidirectional CVIS. Despite the 3D reconstruction ability of GFMs, directly applying them to CVGL and CVIS remains challenging due to the large viewpoint gap between ground and aerial imagery. We propose GeoMap, which embeds ground and aerial features into a shared 3D-aware latent space, effectively reducing cross-view discrepancies for localization. This shared latent space naturally bridges cross-view image synthesis in both directions. To exploit this, we propose GeoFlow, a flow-matching model conditioned on geometry-aware latent embeddings. We further introduce a consistency loss to enforce latent alignment between the two synthesis directions, ensuring bidirectional coherence. Extensive experiments on standard benchmarks, including CVUSA, CVACT, and VIGOR, demonstrate that Geo 2 achieves state-of-the-art performance in both localization and synthesis, highlighting the effectiveness of 3D geometric priors for cross-view geospatial learning. Our source code can be accessed through https://fobow.github.io/geo2.github.io/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper28
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer 等CVPR 2022 · 被引用 6,782 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 被引用 5,687 次
- DUSt3R: Geometric 3D Vision Made EasyShuzhe Wang, Vincent Leroy, Yohann Cabon, Boris Chidlovskii 等CVPR 2024 · 被引用 302 次
相关 Paper
- Aligning Geometric Spatial Layout in Cross-View Geo-Localization via Feature RecombinationQingwang Zhang, Yingying ZhuAAAI 2024 · 被引用 28 次
- Cross-View Geo-Localization via Learning Disentangled Geometric Layout CorrespondenceXiaohan Zhang, Xingyu Li, Waqas Sultani, Yi Zhou 等AAAI 2023 · 被引用 111 次
- Augmenting Cross-View Geo-Localization with Spatial Semantics from Vision Foundation ModelsJi Shen, Lixing Chen, Yang Bai, Zhongqi Miao 等WWW 2026
- Coming Down to Earth: Satellite-to-Street View Synthesis for Geo-LocalizationAysim Toker, Qunjie Zhou, Maxim Maximov, Laura Leal-TaixéCVPR 2021
- CVGL: Causal Learning and Geometric TopologySongsong Ouyang, Yingying ZhuNeurIPS 2025 · 被引用 2 次
