FUSAR-GPT: A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery
Xiaokun Zhang, Yi Yang, Ziqi Ye, Baiyun, Xiaorong Guo, Qingchen Fang, Ruyi Zhang, Xinpeng Zhou, Haipeng Wang
摘要
Research on the intelligent interpretation of allweather, all-time Synthetic Aperture Radar (SAR) is crucial for advancing remote sensing applications. In recent years, although Visual Language Models (VLMs) have demonstrated strong open-world understanding capabilities on RGB images, their performance is severely limited when directly applied to the SAR field due to the complexity of the imaging mechanism, sensitivity to scattering features, and the scarcity of high-quality text corpora. To systematically address this issue, we constructed the inaugural SAR Image-Text-AlphaEarth feature triplet dataset and developed FUSAR-GPT, a VLM specifically for SAR. FUSAR-GPT innovatively introduces a geospatial baseline model as a "world knowledge" prior and embeds multi-source remote-sensing temporal features into the model's visual backbone via "spatiotemporal anchors", enabling dynamic compensation for the sparse representation of targets in SAR images. Furthermore, we designed a two-stage SFT strategy to decouple the knowledge injection and task execution of large models. The spatiotemporal feature embedding and the two-stage decoupling paradigm enable FUSAR-GPT to achieve state-ofthe-art performance across several typical remote sensing visual-language benchmark tests, significantly outperforming mainstream baseline models by over 10%.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- R3Det: Refined Single-Stage Detector with Feature Refinement for Rotating ObjectXue Yang, Junchi Yan, Ziming Feng, Tao HeAAAI 2021 · 被引用 1,109 次
- SkyScript: A Large and Semantically Diverse Vision-Language Dataset for Remote SensingZhecheng Wang, Rajanie Prabha, Tianyuan Huang, Jiajun Wu 等AAAI 2024 · 被引用 167 次
- VHM: Versatile and Honest Vision Language Model for Remote Sensing Image AnalysisChao Pang, Xingxing Weng, Jiang Wu, Jiayu Li 等AAAI 2025 · 被引用 78 次
相关 Paper
- SkySense-VITA: Towards Universal In-context Segmentation of Multi-modal Remote Sensing ImageryKang Wu, Lei Yu, Junwei Luo, Bo Dang 等CVPR 2026 · 被引用 1 次
- GeoChat: Grounded Large Vision-Language Model for Remote SensingKartik Kuckreja, Muhammad Sohail Danish, Muzammal Naseer, Abhijit Das 等CVPR 2024
- T-APT: Text-Guided Modality-Aware Prompt Tuning for Arbitrary Multimodal Remote Sensing Data Joint ClassificationQinghao Gao, Jiahui Qu, Wenqian DongAAAI 2026
- MM-OVSeg: Multimodal Optical-SAR Fusion for Open-Vocabulary Segmentation in Remote SensingYimin Wei, Aoran Xiao, Hongruixuan Chen, Junshi Xia 等CVPR 2026 · 被引用 6 次
- Text Before Vision: Staged Knowledge Injection Matters for Agentic RLVR in Ultra-High-Resolution Remote Sensing UnderstandingFengxiang Wang, Mingshuo Chen, Yueying Li, Yulin Wang 等ICML 2026 · 被引用 1 次
