ERGeoBench: A Comprehensive Benchmark for Embodied Reasoning and Geo-localization in Multimodal Large Language Models
Kaiwen Xue, Tao Wei, Guoxin Zhang, Zhonghong Ou, Kaoyan Lu, Yu Feng, Yifan Zhu, Haoran Luo
摘要
Multimodal large language models (MLLMs) have shown strong potential as embodied agents, yet embodied geo-localization remains underexplored due to the lack of fine-grained evaluation. We introduce ERGeoBench, a diagnostic benchmark for vision-driven embodied geo-localization. ERGeoBench evaluates models under three progressive settings---single-view, panorama-view, and embodied-view---where agents may actively acquire observations through sequential changes in yaw, pitch, and zoom. The benchmark contains 2,207 globally distributed street-view panoramas and measures four complementary capabilities: foundational perception, spatial awareness, common sense reasoning, and geo-localization reasoning. Evaluations of leading proprietary and open-source MLLMs show that current models can infer high-level geographic semantics, but still struggle with fine-grained perceptual operations, metric localization, and spatial consistency across views. We further observe that geo-localization is strongly correlated with the other capability dimensions, suggesting that accurate localization depends on integrated perception, spatial reasoning, and commonsense inference rather than isolated visual recognition. Overall, ERGeoBench provides a unified framework for diagnosing and advancing human-like embodied geo-localization. Project Page: https://kaixuewen.github.io/ERGeoBench/
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- Correlation Verification for Image RetrievalSeongwon Lee, Hongje Seong, Suhyeon Lee, Euntai KimCVPR 2022 · 被引用 79 次
- Global Features are All You Need for Image Retrieval and RerankingShihao Shao, Kaifeng Chen, Arjun Karpur, Qinghua Cui 等ICCV 2023 · 被引用 69 次
- G3: An Effective and Adaptive Framework for Worldwide Geolocalization Using Large Multi-Modality ModelsPengyue Jia, Yiding Liu, Xiaopeng Li, Xiangyu Zhao 等NeurIPS 2024 · 被引用 60 次
- Recognition through Reasoning: Reinforcing Image Geo-localization with Large Vision-Language ModelsLing Li, Yao Zhou, Yuxuan Liang, Fugee Tsung 等NeurIPS 2025 · 被引用 30 次
- GRE Suite: Geo-localization Inference via Fine-Tuned Vision-Language Models and Enhanced Reasoning ChainsChun Wang, Xiaojun Ye, Xiaoran Pan, Zihao Pan 等NeurIPS 2025 · 被引用 18 次
相关 Paper
- EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied AgentsRui Yang, Hanyang Chen, Junyu Zhang, Mark Zhao 等ICML 2025
- GeoX-Bench: Benchmarking Cross-View Geo-Localization and Pose Estimation Capabilities of Large Multimodal ModelsYushuo Zheng, Jiangyong Ying, Huiyu Duan, Chunyi Li 等AAAI 2026 · 被引用 2 次
- Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMsChun-Hsiao Yeh, Chenyu Wang, Shengbang Tong, Ta Ying Cheng 等AAAI 2026 · 被引用 35 次
- UrBench: A Comprehensive Benchmark for Evaluating Large Multimodal Models in Multi-View Urban ScenariosBaichuan Zhou, Haote Yang, Dairong Chen, Junyan Ye 等AAAI 2025 · 被引用 34 次
- UrbanVideo-Bench: Benchmarking Vision-Language Models on Embodied Intelligence with Video Data in Urban SpacesBaining Zhao, Jianjie Fang, Zichao Dai, Ziyou Wang 等ACL 2025 · 被引用 31 次
