SkyScript: A Large and Semantically Diverse Vision-Language Dataset for Remote Sensing
Zhecheng Wang, Rajanie Prabha, Tianyuan Huang, Jiajun Wu, Ram Rajagopal
Abstract
Remote sensing imagery, despite its broad applications in helping achieve Sustainable Development Goals and tackle climate change, has not yet benefited from the recent advancements of versatile, task-agnostic vision language models (VLMs). A key reason is that the large-scale, semantically diverse image-text dataset required for developing VLMs is still absent for remote sensing images. Unlike natural images, remote sensing images and their associated text descriptions cannot be efficiently collected from the public Internet at scale. In this work, we bridge this gap by using geo-coordinates to automatically connect open, unlabeled remote sensing images with rich semantics covered in OpenStreetMap, and thus construct SkyScript, a comprehensive vision-language dataset for remote sensing images, comprising 2.6 million image-text pairs covering 29K distinct semantic tags. With continual pre-training on this dataset, we obtain a VLM that surpasses baseline models with a 6.2% average accuracy gain in zero-shot scene classification across seven benchmark datasets. It also demonstrates the ability of zero-shot transfer for fine-grained object attribute classification and cross-modal retrieval. We hope this dataset can support the advancement of VLMs for various multi-modal tasks in remote sensing, such as open-vocabulary classification, retrieval, captioning, and text-to-image synthesis.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3ed56fb4-afff-41d5-a9c9-27d33230970eCited by top-tier papers29
- VHM: Versatile and Honest Vision Language Model for Remote Sensing Image AnalysisChao Pang, Xingxing Weng, Jiang Wu, Jiayu Li et al.AAAI 2025 · 78 citations
- Earth-Agent: Unlocking the Full Landscape of Earth Observation with AgentsPeilin Feng, Zhutao Lv, Junyan Ye, Xiaolei Wang et al.ICLR 2026 · 49 citations
- TerraFM: A Scalable Foundation Model for Unified Multisensor Earth ObservationMuhammad Sohail Danish, Muhammad Akhtar Munir, Syed Roshaan Ali Shah, Muhammad Haris Khan et al.ICLR 2026 · 30 citations
- Quality-Driven Curation of Remote Sensing Vision-Language Data via Learned Scoring ModelsDilxat Muhtar, Enzhuo Zhang, Zhenshi Li, Feng Gu et al.NeurIPS 2025 · 16 citations
- InstructSAM: A Training-free Framework for Instruction-Oriented Remote Sensing Object RecognitionYijie Zheng, Weijie Wu, Qingyun Li, Xuehui Wang et al.NeurIPS 2025 · 12 citations
Builds on10
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- SatMAE: Pre-training Transformers for Temporal and Multi-Spectral Satellite ImageryYezhen Cong, Samar Khanna, Chenlin Meng, Patrick Liu et al.NeurIPS 2022 · 707 citations
Related papers
- Remote Sensing Vision-Language Foundation Models without Annotations via Ground Remote AlignmentUtkarsh Mall, Cheng Perng Phoo, Meilin Kelsey Liu, Carl Vondrick et al.ICLR 2024 · 90 citations
- GeoChat: Grounded Large Vision-Language Model for Remote SensingKartik Kuckreja, Muhammad Sohail Danish, Muzammal Naseer, Abhijit Das et al.CVPR 2024
- ZoRI: Towards Discriminative Zero-Shot Remote Sensing Instance SegmentationShiqi Huang, Shuting He, Bihan WenAAAI 2025
- Towards Open-Vocabulary Remote Sensing Image Semantic SegmentationChengyang Ye, Yunzhi Zhuge, Pingping ZhangAAAI 2025 · 31 citations
- Any2RSI: Controllable Remote Sensing Text-to-Image Generation via Any Control and Enriched DescriptionXu Zhang, Jianzhong Huang, Lefei ZhangAAAI 2026 · 1 citation
