GeoArena: Evaluating Open-World Geographic Reasoning in Large Vision-Language Models
Pengyue Jia, Yingyi Zhang, Xiangyu Zhao, Sharon Li
Abstract
Geographic reasoning is a fundamental cognitive capability that requires models to infer plausible locations by synthesizing visual evidence with spatial world knowledge. Despite recent advances in large vision-language models (LVLMs), existing evaluation paradigms remain largely outcome-centric, relying on static datasets and predefined labels that are conceptually misaligned with open-world geographic inference. Such outcome-centric evaluations often focus exclusively on label matching, leaving the underlying linguistic reasoning chains as unexamined black boxes. In this work, we introduce GeoArena, a dynamic, human-preference-based evaluation framework for benchmarking open-world geographic reasoning. GeoArena reframes evaluation as a pairwise reasoning alignment task on in-thewild images, where human judges compare model-generated explanations based on reasoning quality, evidence synthesis, and plausibility. We deploy GeoArena as a public platform and benchmark 17 frontier LVLMs using thousands of human judgments, which complements existing benchmarks and supports the development of geographically grounded, humanaligned AI systems. We further provide detailed analyses of model behavior, including reliability of human preferences and factors influencing judgments of geographic reasoning quality. We open-source GeoArena 1 to foster future research.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cf9290ed-3db3-48ef-ad03-27c1b43832cbBuilds on14
- GeoCLIP: Clip-Inspired Alignment between Locations and Images for Effective Worldwide Geo-localizationVicente Vivanco Cepeda, Gaurav Kumar Nayak, Mubarak ShahNeurIPS 2023 · 303 citations
- Multi-Type Urban Crime PredictionXiangyu Zhao, Wenqi Fan, Hui Liu, Jiliang TangAAAI 2022 · 36 citations
- GeoReasoner: Geo-localization with Reasoning in Street Views using a Large Vision-Language ModelLing Li, Yu Ye, Bingchuan Jiang, Wei ZengICML 2024 · 35 citations
- AutoSTL: Automated Spatio-Temporal Multi-Task LearningZijian Zhang, Xiangyu Zhao, Hao Miao, Chunxu Zhang et al.AAAI 2023 · 31 citations
- Recognition through Reasoning: Reinforcing Image Geo-localization with Large Vision-Language ModelsLing Li, Yao Zhou, Yuxuan Liang, Fugee Tsung et al.NeurIPS 2025 · 30 citations
Related papers
- UrbanGeoEval: A City-Scale Benchmark for Evaluating Large Language Models in Geospatial ReasoningMutian Bao, Qiuyi Qi, Tian Liang, Jinjian Zhang et al.ACL 2026
- GeoRC: A Benchmark for Geolocation Reasoning ChainsMohit Talreja, Joshua Diao, Jim James, Radu Casapu et al.ACL 2026 · 1 citation
- GameArena: Evaluating LLM Reasoning through Live Computer GamesLanxiang Hu, Qiyu Li, Anze Xie, Nan Jiang et al.ICLR 2025
- TimeSpot: Benchmarking Geo-Temporal Understanding in Vision–Language Models in Real-World SettingsAzmine Toushik Wasi, Shahriyar Zaman Ridoy, Koushik Ahamed Tonmoy, Kinga Tshering et al.ICML 2026
- GRE Suite: Geo-localization Inference via Fine-Tuned Vision-Language Models and Enhanced Reasoning ChainsChun Wang, Xiaojun Ye, Xiaoran Pan, Zihao Pan et al.NeurIPS 2025 · 18 citations
