GEOBench-VLM: Benchmarking Vision-Language Models for Geospatial Tasks
Muhammad Sohail Danish, Muhammad Akhtar Munir, Syed Roshaan Ali Shah, Kartik Kuckreja, Fahad Shahbaz Khan, Paolo Fraccaro, Alexandre Lacoste, Salman Khan
Abstract
Describe the spatial relationship between object in green box and object in red box. A. A small vehicle is to the left of the baseballdiamond B. A small vehicle is to the right of the large-vehicle C. A ship is below the harbor D. A large vehicle is beside the storage-tank E. A small vehicle is above the bridge. Q: Describe the spatial relationship between object in green box and object in red box. A. A small vehicle is to the left of the baseballdiamond B. A small vehicle is to the right of the large-vehicle C. A ship is below the harbor D. A large vehicle is beside the storage-tank E. A small vehicle is above the bridge. Q: Describe the spatial relationship between object in green box and object in red box. A. A small vehicle is to the left of the baseballdiamond B. A small vehicle is to the right of the large-vehicle C. A ship is below the harbor D. A large vehicle is beside the storage-tank E. A small vehicle is above the bridge. Q: Describe the image in detail. Cap: The aerial image showcases an industrial area adjacent to a large harbor. In the top section, a series of white storage tanks are neatly aligned, forming a structured pattern. These tanks are positioned primarily in the top … Q: Describe the image in detail. Cap: The aerial image showcases an industrial area adjacent to a large harbor. In the top section, a series of white storage tanks are neatly aligned, forming a structured pattern. These tanks are positioned primarily in the top … Q: Describe the image in detail. Cap: The aerial image showcases an industrial area adjacent to a large harbor. In the top section, a series of white storage tanks are neatly aligned, forming a structured pattern. These tanks are positioned primarily in the top … Q: What is the estimated magnitude of the earthquake that affected this area? A.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c19e193e-6b0e-4da5-8e57-593295fcb792Cited by top-tier papers11
- OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language ModelsMengdi Jia, Zekun Qi, Shaochen Zhang, Wenyao Zhang et al.ICLR 2026 · 109 citations
- SegEarth-R2: Towards Comprehensive Language-guided Segmentation for Remote Sensing ImagesZepeng Xin, Kaiyu Li, Luodi Chen, Wanchen Li et al.CVPR 2026 · 14 citations
- TerraScope: Pixel-Grounded Visual Reasoning for Earth ObservationYan Shu, Bin Ren, Zhitong Xiong, Xiao Xiang Zhu et al.CVPR 2026 · 9 citations
- GeoMMBench and GeoMMAgent: Toward Expert-Level Multimodal Intelligence in Geoscience and Remote SensingAoran Xiao, Shihao Cheng, Yonghao Xu, Yexian Ren et al.CVPR 2026 · 6 citations
- GeoArena: Evaluating Open-World Geographic Reasoning in Large Vision-Language ModelsPengyue Jia, Yingyi Zhang, Xiangyu Zhao, Sharon LiACL 2026 · 3 citations
Builds on21
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- MM-Vet: Evaluating Large Multimodal Models for Integrated CapabilitiesWeihao Yu, Zhengyuan Yang, Linjie Li, Jianfeng Wang et al.ICML 2024 · 1,191 citations
Related papers
- Action-Sketcher: From Reasoning to Action via Visual Sketches for Robotic ManipulationHuajie Tan, Peterson Co, Yijie Xu, Shanyu Rong et al.CVPR 2026
- Argus: Vision-Centric Reasoning with Grounded Chain-of-ThoughtYunze Man, De-An Huang, Guilin Liu, Shiwei Sheng et al.CVPR 2025
- Auto-Controlled Image Perception in MLLMs via Visual Perception TokensRunpeng Yu, Xinyin Ma, Xinchao WangICCV 2025 · 1 citation
- Towards Backward-Compatible Continual Learning of Image CompressionZhihao Duan, Ming Lu, Justin Yang, Jiangpeng He et al.CVPR 2024
- Multiagent Multitraversal Multimodal Self-Driving: Open MARS DatasetYiming Li, Zhiheng Li, Nuo Chen, Moonjun Gong et al.CVPR 2024
