GeoSURGE: Geo-localization using Semantic Fusion with Hierarchy of Geographic Embeddings
Angel Daruna, Nicholas Meegan, Han-Pang Chiu, Supun Samarasekera, Rakesh Kumar
Abstract
Worldwide visual geo-localization aims to determine the geographic location of an image anywhere on Earth using only its visual content. Despite recent progress, learning expressive representations of geographic space remains challenging due to the inherently low-dimensional nature of geographic coordinates. We formulate global geo-localization as aligning the visual representation of a query image with a learned geographic representation. Our approach explicitly models the world as a hierarchy of learned geographic embeddings, enabling a distributed and multi-scale representation of geographic space. In addition, we introduce a semantic fusion module that efficiently integrates appearance features with semantic segmentation through latent cross-attention, producing a more robust visual representation for localization. Experiments on five widely used geo-localization benchmarks demonstrate that our method achieves new state-of-the-art results on 22 of 25 reported metrics. Ablation studies show that these improvements are primarily driven by the proposed geographic representation and semantic fusion mechanism.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5305b0df-568e-4e33-b6db-ece3f9a300a0Builds on9
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- GeoCLIP: Clip-Inspired Alignment between Locations and Images for Effective Worldwide Geo-localizationVicente Vivanco Cepeda, Gaurav Kumar Nayak, Mubarak ShahNeurIPS 2023 · 303 citations
- G3: An Effective and Adaptive Framework for Worldwide Geolocalization Using Large Multi-Modality ModelsPengyue Jia, Yiding Liu, Xiaopeng Li, Xiangyu Zhao et al.NeurIPS 2024 · 60 citations
- GeoReasoner: Geo-localization with Reasoning in Street Views using a Large Vision-Language ModelLing Li, Yu Ye, Bingchuan Jiang, Wei ZengICML 2024 · 35 citations
- OpenStreetView-5M: The Many Roads to Global Visual GeolocationGuillaume Astruc, Nicolas Dufour, Ioannis Siglidis, Constantin Aronssohn et al.CVPR 2024
Related papers
- Where We Are and What We're Looking At: Query Based Worldwide Image Geo-localization Using Hierarchies and ScenesBrandon Clark, Alec Kerrigan, Parth Parag Kulkarni, Vicente Vivanco Cepeda et al.CVPR 2023
- Scaling Image Geo-Localization to Continent LevelPhilipp Lindenberger, Paul-Edouard Sarlin, Jan Hosang, Marc Pollefeys et al.NeurIPS 2025 · 11 citations
- Cross-View Visual Geo-Localization for Outdoor Augmented RealityNiluthpol Chowdhury Mithun, Kshitij Minhas, Han-Pang Chiu, Taragay Oskiper et al.IEEE VR 2023 · 22 citations
- VIGOR: Cross-View Image Geo-Localization Beyond One-to-One RetrievalSijie Zhu, Taojiannan Yang, Chen ChenCVPR 2021
- Coming Down to Earth: Satellite-to-Street View Synthesis for Geo-LocalizationAysim Toker, Qunjie Zhou, Maxim Maximov, Laura Leal-TaixéCVPR 2021
