GEOBench-VLM: Benchmarking Vision-Language Models for Geospatial Tasks
Muhammad Sohail Danish, Muhammad Akhtar Munir, Syed Roshaan Ali Shah, Kartik Kuckreja, Fahad Shahbaz Khan, Paolo Fraccaro, Alexandre Lacoste, Salman Khan
摘要
Describe the spatial relationship between object in green box and object in red box. A. A small vehicle is to the left of the baseballdiamond B. A small vehicle is to the right of the large-vehicle C. A ship is below the harbor D. A large vehicle is beside the storage-tank E. A small vehicle is above the bridge. Q: Describe the spatial relationship between object in green box and object in red box. A. A small vehicle is to the left of the baseballdiamond B. A small vehicle is to the right of the large-vehicle C. A ship is below the harbor D. A large vehicle is beside the storage-tank E. A small vehicle is above the bridge. Q: Describe the spatial relationship between object in green box and object in red box. A. A small vehicle is to the left of the baseballdiamond B. A small vehicle is to the right of the large-vehicle C. A ship is below the harbor D. A large vehicle is beside the storage-tank E. A small vehicle is above the bridge. Q: Describe the image in detail. Cap: The aerial image showcases an industrial area adjacent to a large harbor. In the top section, a series of white storage tanks are neatly aligned, forming a structured pattern. These tanks are positioned primarily in the top … Q: Describe the image in detail. Cap: The aerial image showcases an industrial area adjacent to a large harbor. In the top section, a series of white storage tanks are neatly aligned, forming a structured pattern. These tanks are positioned primarily in the top … Q: Describe the image in detail. Cap: The aerial image showcases an industrial area adjacent to a large harbor. In the top section, a series of white storage tanks are neatly aligned, forming a structured pattern. These tanks are positioned primarily in the top … Q: What is the estimated magnitude of the earthquake that affected this area? A.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language ModelsMengdi Jia, Zekun Qi, Shaochen Zhang, Wenyao Zhang 等ICLR 2026 · 被引用 109 次
- SegEarth-R2: Towards Comprehensive Language-guided Segmentation for Remote Sensing ImagesZepeng Xin, Kaiyu Li, Luodi Chen, Wanchen Li 等CVPR 2026 · 被引用 14 次
- TerraScope: Pixel-Grounded Visual Reasoning for Earth ObservationYan Shu, Bin Ren, Zhitong Xiong, Xiao Xiang Zhu 等CVPR 2026 · 被引用 9 次
- GeoMMBench and GeoMMAgent: Toward Expert-Level Multimodal Intelligence in Geoscience and Remote SensingAoran Xiao, Shihao Cheng, Yonghao Xu, Yexian Ren 等CVPR 2026 · 被引用 6 次
- GeoArena: Evaluating Open-World Geographic Reasoning in Large Vision-Language ModelsPengyue Jia, Yingyi Zhang, Xiangyu Zhao, Sharon LiACL 2026 · 被引用 3 次
它引用的顶会 Paper21
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar 等NeurIPS 2021 · 被引用 9,661 次
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- MM-Vet: Evaluating Large Multimodal Models for Integrated CapabilitiesWeihao Yu, Zhengyuan Yang, Linjie Li, Jianfeng Wang 等ICML 2024 · 被引用 1,191 次
相关 Paper
- Action-Sketcher: From Reasoning to Action via Visual Sketches for Robotic ManipulationHuajie Tan, Peterson Co, Yijie Xu, Shanyu Rong 等CVPR 2026
- Argus: Vision-Centric Reasoning with Grounded Chain-of-ThoughtYunze Man, De-An Huang, Guilin Liu, Shiwei Sheng 等CVPR 2025
- Auto-Controlled Image Perception in MLLMs via Visual Perception TokensRunpeng Yu, Xinyin Ma, Xinchao WangICCV 2025 · 被引用 1 次
- Towards Backward-Compatible Continual Learning of Image CompressionZhihao Duan, Ming Lu, Justin Yang, Jiangpeng He 等CVPR 2024
- Multiagent Multitraversal Multimodal Self-Driving: Open MARS DatasetYiming Li, Zhiheng Li, Nuo Chen, Moonjun Gong 等CVPR 2024
