UrbanGeoEval: A City-Scale Benchmark for Evaluating Large Language Models in Geospatial Reasoning
Mutian Bao, Qiuyi Qi, Tian Liang, Jinjian Zhang, Wei Zhou, Ming Kong, Linjian Mo, Qiang Zhu
摘要
Current evaluations of geospatial reasoning in LLMs are frequently impeded by the entanglement of factual recall and spatial logic, which often obscures the models' true capabilities in complex city-scale environments. To address this, we introduce UrbanGeoEval, a comprehensive benchmark featuring a dual-module framework designed to disentangle these competencies. The Knowledge Module assesses urban memory via scalable map-based queries, while the Reasoning Module isolates pure logical inference across 3,148 realistic tasks by providing necessary geospatial context. Unlike prior benchmarks that hand the model pre-computed spatial text, UrbanGeoEval provides raw geometry and forces the model to act as a spatial computing engine. Our evaluation methodology introduces a reliable hybrid pipeline that merges deterministic programmatic checks with an LLM-as-a-Judge, achieving expert-level evaluation accuracy. Extensive experiments on 18 widely used LLMs uncover critical insights: (1) models exhibit severe geographic biases and resolution gaps; (2) failures in complex multi-hop tasks often stem from brittle foundational spatial skills rather than high-level logic deficits. UrbanGeoEval provides a precise diagnostic tool for advancing urban geospatial intelligence in LLMs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- Urban2Vec: Incorporating Street View Imagery and POIs for Multi-Modal Urban Neighborhood EmbeddingZhecheng Wang, Haoyuan Li, Ram RajagopalAAAI 2020 · 被引用 113 次
- GeoLLM: Extracting Geospatial Knowledge from Large Language ModelsRohin Manvi, Samar Khanna, Gengchen Mai, Marshall Burke 等ICLR 2024 · 被引用 104 次
- MGeo: Multi-Modal Geographic Language Model Pre-TrainingRuixue Ding, Boli Chen, Pengjun Xie, Fei Huang 等SIGIR 2023 · 被引用 29 次
- CityGPT: Empowering Urban Spatial Cognition of Large Language ModelsJie Feng, Tianhui Liu, Yuwei Du, Siqi Guo 等KDD 2025 · 被引用 9 次
相关 Paper
- UrBench: A Comprehensive Benchmark for Evaluating Large Multimodal Models in Multi-View Urban ScenariosBaichuan Zhou, Haote Yang, Dairong Chen, Junyan Ye 等AAAI 2025 · 被引用 34 次
- GGBench: A Geometric Generative Reasoning Benchmark for Unified Multimodal ModelsJingxuan Wei, Caijun Jia, Xi Bai, Xinglong Xu 等CVPR 2026 · 被引用 7 次
- USTBench: Benchmarking and Dissecting Spatiotemporal Reasoning Capabilities of LLMs as Urban AgentsSiqi Lai, Yansong Ning, Zirui Yuan, Zhixi Chen 等ICLR 2026 · 被引用 7 次
- Is A Picture Worth A Thousand Words? Delving Into Spatial Reasoning for Vision Language ModelsJiayu Wang, Yifei Ming, Zhenmei Shi, Vibhav Vineet 等NeurIPS 2024 · 被引用 166 次
- MapEval: A Map-Based Evaluation of Geo-Spatial Reasoning in Foundation ModelsMahir Labib Dihan, Md Tanvir Hassan, Md Tanvir Parvez, Md Hasebul Hasan 等ICML 2025
