UrBench: A Comprehensive Benchmark for Evaluating Large Multimodal Models in Multi-View Urban Scenarios
Baichuan Zhou, Haote Yang, Dairong Chen, Junyan Ye, Tianyi Bai, Jinhua Yu, Songyang Zhang, Dahua Lin, Conghui He, Weijia Li
摘要
Recent evaluations of Large Multimodal Models (LMMs) have explored their capabilities in various domains, with only few benchmarks specifically focusing on urban environments. Moreover, existing urban benchmarks have been limited to evaluating LMMs with basic region-level urban tasks under singular views, leading to incomplete evaluations of LMMs' abilities in urban environments. To address these issues, we present UrBench, a comprehensive benchmark designed for evaluating LMMs in complex multi-view urban scenarios. UrBench contains 11.6K meticulously curated questions at both region-level and role-level that cover 4 task dimensions: Geo-Localization, Scene Reasoning, Scene Understanding, and Object Understanding, totaling 14 task types. In constructing UrBench, we utilize data from existing datasets and additionally collect data from 11 cities, creating new annotations using a cross-view detection-matching method. With these images and annotations, we then integrate LMM-based, rule-based, and human-based methods to construct large-scale high-quality questions. Our evaluations on 21 LMMs show that current LMMs struggle in the urban environments in several aspects. Even the best performing GPT-4o lags behind humans in most tasks, ranging from simple tasks such as counting to complex tasks such as orientation, localization and object attribute recognition, with an average performance gap of 17.4%. Our benchmark also reveals that LMMs exhibit inconsistent behaviors with different urban views, especially with respect to understanding cross-view relations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Spot the Fake: Large Multimodal Model-Based Synthetic Image Detection with Artifact ExplanationSiwei Wen, Junyan Ye, Peilin Feng, Hengrui Kang 等NeurIPS 2025 · 被引用 82 次
- Earth-Agent: Unlocking the Full Landscape of Earth Observation with AgentsPeilin Feng, Zhutao Lv, Junyan Ye, Xiaolei Wang 等ICLR 2026 · 被引用 49 次
- CityLens: Evaluating Large Vision-Language Models for Urban Socioeconomic SensingTianhui Liu, Hetian Pang, Xin Zhang, Tianjian Ouyang 等ICLR 2026 · 被引用 10 次
- AirCopBench: A Benchmark for Multi-drone Collaborative Embodied Perception and ReasoningJirong Zha, Yuxuan Fan, Tianyu Zhang, Geng Chen 等AAAI 2026 · 被引用 9 次
- UrbanLLaVA: A Multi-Modal Large Language Model for Urban IntelligenceJie Feng, Shengyuan Wang, Tianhui Liu, Yanxin Xi 等ICCV 2025 · 被引用 7 次
它引用的顶会 Paper13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- MMMU: A Massive Multi-Discipline Multimodal Understanding and Reasoning Benchmark for Expert AGIXiang Yue, Yuansheng Ni, Tianyu Zheng, Kai Zhang 等CVPR 2024 · 被引用 213 次
- UrbanCLIP: Learning Text-enhanced Urban Region Profiling with Contrastive Language-Image Pretraining from the WebYibo Yan, Haomin Wen, Siru Zhong, Wei Chen 等WWW 2024 · 被引用 124 次
- VELMA: Verbalization Embodiment of LLM Agents for Vision and Language Navigation in Street ViewRaphael Schumann, Wanrong Zhu, Weixi Feng, Tsu-Jui Fu 等AAAI 2024 · 被引用 122 次
相关 Paper
- Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMsChun-Hsiao Yeh, Chenyu Wang, Shengbang Tong, Ta Ying Cheng 等AAAI 2026 · 被引用 35 次
- ERGeoBench: A Comprehensive Benchmark for Embodied Reasoning and Geo-localization in Multimodal Large Language ModelsKaiwen Xue, Tao Wei, Guoxin Zhang, Zhonghong Ou 等ICML 2026
- 4D-Bench: Benchmarking Multi-Modal Large Language Models for 4D Object UnderstandingWenxuan Zhu, Bing Li, Cheng Zheng, Jinjie Mai 等ICCV 2025 · 被引用 2 次
- USTBench: Benchmarking and Dissecting Spatiotemporal Reasoning Capabilities of LLMs as Urban AgentsSiqi Lai, Yansong Ning, Zirui Yuan, Zhixi Chen 等ICLR 2026 · 被引用 7 次
- OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMsCaorui Li, Yu Chen, Yiyan Ji, Jin Xu 等ICLR 2026 · 被引用 53 次
