GRE Suite: Geo-localization Inference via Fine-Tuned Vision-Language Models and Enhanced Reasoning Chains
Chun Wang, Xiaojun Ye, Xiaoran Pan, Zihao Pan, Haofan Wang, Yiren Song
Abstract
Recent advances in Visual Language Models (VLMs) have demonstrated exceptional performance in visual reasoning tasks. However, geo-localization presents unique challenges, requiring the extraction of multigranular visual cues from images and their integration with external world knowledge for systematic reasoning. Current approaches to geo-localization tasks often lack robust reasoning mechanisms and explainability, limiting their effectiveness. To address these limitations, we propose the Geo Reason Enhancement (GRE) Suite, a novel framework that augments VLMs with structured reasoning chains for accurate and interpretable location inference. The GRE Suite is systematically developed across three key dimensions: dataset, model, and benchmark. First, we introduce GRE30K, a high-quality geo-localization reasoning dataset designed to facilitate fine-grained visual and contextual analysis. Next, we present the GRE model, which employs a multi-stage reasoning strategy to progressively infer scene attributes, local details, and semantic features, thereby narrowing down potential geographic regions with enhanced precision. Finally, we construct the Geo Reason Evaluation Benchmark (GREval-Bench), a comprehensive evaluation framework that assesses VLMs across diverse urban, natural, and landmark scenes to measure both coarse-grained (e.g., country, continent) and fine-grained (e.g., city, street) localization performance. Experimental results demonstrate that GRE significantly outperforms existing methods across all granularities of geo-localization tasks, underscoring the efficacy of reasoning-augmented VLMs in complex geographic inference. Code and data will be released at https://github.com/Thorin215/GRE.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 18675834-178c-42e5-9107-27d76e1fcfa7Cited by top-tier papers4
- GeoAgent: Learning to Geolocate Everywhere with Reinforced Geographic CharacteristicsModi Jin, Yiming Zhang, Boyuan Sun, Dingwen Zhang et al.CVPR 2026 · 7 citations
- SpotAgent: Grounding Visual Geo-localization in Large Vision-Language Models through Agentic ReasoningFurong Jia, Ling Dai, Wenjin Deng, Fan Zhang et al.KDD 2026 · 6 citations
- GeoRC: A Benchmark for Geolocation Reasoning ChainsMohit Talreja, Joshua Diao, Jim James, Radu Casapu et al.ACL 2026 · 1 citation
- ERGeoBench: A Comprehensive Benchmark for Embodied Reasoning and Geo-localization in Multimodal Large Language ModelsKaiwen Xue, Tao Wei, Guoxin Zhang, Zhonghong Ou et al.ICML 2026
Builds on40
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- CLIPScore: A Reference-free Evaluation Metric for Image CaptioningJack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras et al.EMNLP 2021 · 937 citations
Related papers
- Recognition through Reasoning: Reinforcing Image Geo-localization with Large Vision-Language ModelsLing Li, Yao Zhou, Yuxuan Liang, Fugee Tsung et al.NeurIPS 2025 · 30 citations
- TimeSpot: Benchmarking Geo-Temporal Understanding in Vision–Language Models in Real-World SettingsAzmine Toushik Wasi, Shahriyar Zaman Ridoy, Koushik Ahamed Tonmoy, Kinga Tshering et al.ICML 2026
- GeoReasoner: Geo-localization with Reasoning in Street Views using a Large Vision-Language ModelLing Li, Yu Ye, Bingchuan Jiang, Wei ZengICML 2024 · 35 citations
- GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language FieldsShunsuke Yasuki, Taiki Miyanishi, Nakamasa Inoue, Shuhei Kurita et al.ICCV 2025 · 1 citation
- VisRes Bench: On Evaluating the Visual Reasoning Capabilities of VLMsBrigitta Malagurski Törtei, Yasser Dahou, Ngoc Dung Huynh, Wamiq Reyaz Para et al.CVPR 2026 · 3 citations
