GeoRanker: Distance-Aware Ranking for Worldwide Image Geolocalization
Pengyue Jia, Seongheon Park, Song Gao, Xiangyu Zhao, Sharon Li
Abstract
Worldwide image geolocalization-the task of predicting GPS coordinates from images taken anywhere on Earth-poses a fundamental challenge due to the vast diversity in visual content across regions. While recent approaches adopt a twostage pipeline of retrieving candidates and selecting the best match, they typically rely on simplistic similarity heuristics and point-wise supervision, failing to model spatial relationships among candidates. In this paper, we propose GeoRanker, a distance-aware ranking framework that leverages large vision-language models to jointly encode query-candidate interactions and predict geographic proximity. In addition, we introduce a multi-order distance loss that ranks both absolute and relative distances, enabling the model to reason over structured spatial relationships. To support this, we curate GeoRanking, the first dataset explicitly designed for geographic ranking tasks with multimodal candidate information. GeoRanker achieves state-of-the-art results on two well-established benchmarks (IM2GPS3K and YFCC4K), significantly outperforming current best methods. We also release our code, checkpoint, and dataset online 2 for ease of reproduction.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Recognition through Reasoning: Reinforcing Image Geo-localization with Large Vision-Language ModelsLing Li, Yao Zhou, Yuxuan Liang, Fugee Tsung et al.NeurIPS 2025 · 30 citations
- GeoAgent: Learning to Geolocate Everywhere with Reinforced Geographic CharacteristicsModi Jin, Yiming Zhang, Boyuan Sun, Dingwen Zhang et al.CVPR 2026 · 7 citations
- SpotAgent: Grounding Visual Geo-localization in Large Vision-Language Models through Agentic ReasoningFurong Jia, Ling Dai, Wenjin Deng, Fan Zhang et al.KDD 2026 · 6 citations
- GeoArena: Evaluating Open-World Geographic Reasoning in Large Vision-Language ModelsPengyue Jia, Yingyi Zhang, Xiangyu Zhao, Sharon LiACL 2026 · 3 citations
- VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global ScaleParth Parag Kulkarni, Rohit Gupta, Prakash Chandra Chhipa, Mubarak ShahCVPR 2026 · 1 citation
Builds on19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- GeoCLIP: Clip-Inspired Alignment between Locations and Images for Effective Worldwide Geo-localizationVicente Vivanco Cepeda, Gaurav Kumar Nayak, Mubarak ShahNeurIPS 2023 · 303 citations
- Cross-view Geo-localization with Layer-to-Layer TransformerHongji Yang, Xiufan Lu, Yingying ZhuNeurIPS 2021 · 231 citations
- Instance-level Image Retrieval using Reranking TransformersFuwen Tan, Jiangbo Yuan, Vicente OrdonezICCV 2021 · 116 citations
Related papers
- G3: An Effective and Adaptive Framework for Worldwide Geolocalization Using Large Multi-Modality ModelsPengyue Jia, Yiding Liu, Xiaopeng Li, Xiangyu Zhao et al.NeurIPS 2024 · 60 citations
- VIGOR: Cross-View Image Geo-Localization Beyond One-to-One RetrievalSijie Zhu, Taojiannan Yang, Chen ChenCVPR 2021
- GeoSURGE: Geo-localization using Semantic Fusion with Hierarchy of Geographic EmbeddingsAngel Daruna, Nicholas Meegan, Han-Pang Chiu, Supun Samarasekera et al.CVPR 2026 · 2 citations
- Vision-Language Reasoning for Geolocalization: A Reinforcement Learning ApproachBiao Wu, Meng Fang, Ling Chen, Ke Xu et al.AAAI 2026 · 2 citations
- Cross-View Visual Geo-Localization for Outdoor Augmented RealityNiluthpol Chowdhury Mithun, Kshitij Minhas, Han-Pang Chiu, Taragay Oskiper et al.IEEE VR 2023 · 22 citations
