Towards Faithful Reasoning in Remote Sensing: A Perceptually-Grounded GeoSpatial Chain-of-Thought for Vision-Language Models
Jiaqi Liu, Lang Sun, Ronghao Fu, Bo Yang
Abstract
Vision-Language Models (VLMs) in remote sensing often fail at complex analytical tasks, a limitation stemming from their end-to-end training paradigm that bypasses crucial reasoning steps and leads to unverifiable outputs. To address this limitation, we introduce the Perceptually-Grounded Geospatial Chain-of-Thought (Geo-CoT), a framework that models remote sensing analysis as a verifiable, multi-step process. We instill this analytical process through a two-stage alignment strategy, leveraging Geo-CoT380k, the first large-scale dataset of structured Geo-CoT rationales. This strategy first employs supervised fine-tuning (SFT) to instill the foundational cognitive architecture, then leverages Group Reward Policy Optimization (GRPO) to refine the model's reasoning policy towards factual correctness. The resulting model, RSThinker, outputs both a final answer and its justifying, verifiable analytical trace. This capability yields dominant performance, significantly outperforming state-of-the-art models across a comprehensive range of tasks. The public release of our Geo-CoT380k dataset and RS-Thinker model upon publication serves as a concrete pathway from opaque perception towards structured, verifiable reasoning for Earth Observation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 12e20339-7e96-474a-b50d-3300087ea504Cited by top-tier papers1
Ask how each one uses itBuilds on14
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
- VHM: Versatile and Honest Vision Language Model for Remote Sensing Image AnalysisChao Pang, Xingxing Weng, Jiang Wu, Jiayu Li et al.AAAI 2025 · 78 citations
- Compositional Chain-of-Thought Prompting for Large Multimodal ModelsChancharik Mitra, Brandon Huang, Trevor Darrell, Roei HerzigCVPR 2024 · 62 citations
- MedCoT: Medical Chain of Thought via Hierarchical ExpertJiaxiang Liu, Yuan Wang, Jiawei Du, Joey Zhou et al.EMNLP 2024 · 15 citations
- SkyMoE: A Vision-Language Foundation Model for Enhancing Geospatial Interpretation with Mixture of ExpertsJiaqi Liu, Ronghao Fu, Lang Sun, Haoran Liu et al.AAAI 2026 · 6 citations
Related papers
- GeoCoT: Towards Reliable Remote Sensing Reasoning with Manifold PerspectiveDaixun Li, Zirui Li, Sibo He, Jiayun Tian et al.CVPR 2026
- ReWatch-R1: Boosting Complex Video Reasoning in Large Vision-Language Models through Agentic Data SynthesisCongzhi Zhang, Zhibin Wang, Yinchao Ma, Jiawei Peng et al.ICLR 2026 · 24 citations
- EgoThinker: Unveiling Egocentric Reasoning with Spatio-Temporal CoTBaoqi Pei, Yifei Huang, Jilan Xu, Yuping He et al.NeurIPS 2025 · 21 citations
- Rex-Thinker: Grounded Object Referring via Chain-of-Thought ReasoningQing Jiang, Xingyu Chen, Zhaoyang Zeng, Junzhi Yu et al.ICLR 2026 · 25 citations
- Reason-RFT: Reinforcement Fine-Tuning for Visual Reasoning of Vision Language ModelsHuajie Tan, Yuheng Ji, Xiaoshuai Hao, Xiansheng Chen et al.NeurIPS 2025 · 45 citations
