EarthVQA: Towards Queryable Earth via Relational Reasoning-Based Remote Sensing Visual Question Answering
Junjue Wang, Zhuo Zheng, Zihang Chen, Ailong Ma, Yanfei Zhong
Abstract
Earth vision research typically focuses on extracting geospatial object locations and categories but neglects the exploration of relations between objects and comprehensive reasoning. Based on city planning needs, we develop a multi-modal multi-task VQA dataset (EarthVQA) to advance relational reasoning-based judging, counting, and comprehensive analysis. The EarthVQA dataset contains 6000 images, corresponding semantic masks, and 208,593 QA pairs with urban and rural governance requirements embedded. As objects are the basis for complex relational reasoning, we propose a Semantic OBject Awareness framework (SOBA) to advance VQA in an object-centric way. To preserve refined spatial locations and semantics, SOBA leverages a segmentation network for object semantics generation. The object-guided attention aggregates object interior features via pseudo masks, and bidirectional cross-attention further models object external relations hierarchically. To optimize object counting, we propose a numerical difference loss that dynamically adds difference penalties, unifying the classification and regression tasks. Experimental results show that SOBA outperforms both advanced general and remote sensing methods. We believe this dataset and framework provide a strong benchmark for Earth vision's complex analysis. The project page is at https://Junjue-Wang.github.io/homepage/EarthVQA.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cc4893a4-76ca-4182-8d38-85ff3cee080dCited by top-tier papers15
- Earth-Agent: Unlocking the Full Landscape of Earth Observation with AgentsPeilin Feng, Zhutao Lv, Junyan Ye, Xiaolei Wang et al.ICLR 2026 · 49 citations
- UrBench: A Comprehensive Benchmark for Evaluating Large Multimodal Models in Multi-View Urban ScenariosBaichuan Zhou, Haote Yang, Dairong Chen, Junyan Ye et al.AAAI 2025 · 34 citations
- ZoomEarth: Active Perception for Ultra-High-Resolution Geospatial Vision-Language TasksRuixun Liu, Bowen Fu, Jiayi Song, Kaiyu Li et al.CVPR 2026 · 19 citations
- RemoteReasoner: Towards Unifying Geospatial Reasoning WorkflowLiang Yao, Fan Liu, Hongbo Lu, Chuanyi Zhang et al.AAAI 2026 · 16 citations
- SegEarth-R2: Towards Comprehensive Language-guided Segmentation for Remote Sensing ImagesZepeng Xin, Kaiyu Li, Luodi Chen, Wanchen Li et al.CVPR 2026 · 14 citations
Builds on8
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
Related papers
- Answering Complex Geographic Questions by Adaptive Reasoning with Visual Context and External Commonsense KnowledgeFan Li, Jianxing Yu, Jielong Tang, Wenqing Chen et al.ACL 2025 · 3 citations
- Motion-Grounded Video Reasoning: Understanding and Perceiving Motion at Pixel LevelAndong Deng, Tongjia Chen, Shoubin Yu, Taojiannan Yang et al.CVPR 2025
- VTQA: Visual Text Question Answering via Entity Alignment and Cross-Media ReasoningKang Chen, Xiangqian WuCVPR 2024
- IntentQA: Context-aware Video Intent ReasoningJiapeng Li, Ping Wei, Wenjuan Han, Lifeng FanICCV 2023 · 97 citations
- Re-Attention for Visual Question AnsweringWenya Guo, Ying Zhang, Xiaoping Wu, Jufeng Yang et al.AAAI 2020 · 90 citations
