SkySense-O: Towards Open-World Remote Sensing Interpretation with Vision-Centric Visual-Language Modeling
Qi Zhu, Jiangwei Lao, Deyi Ji, Junwei Luo, Kang Wu, Yingying Zhang, Lixiang Ru, Jian Wang, Jingdong Chen, Ming Yang, Dong Liu, Feng Zhao
Abstract
text alignment, reducing the degradation of general visual representation capabilities of existing paradigms. Then, we construct a visual-relevance knowledge graph across opencategory texts and further develop a novel vision-centric image-text contrastive loss for fine-tuning with text prompts. This new model, denoted as SkySense-O, demonstrates impressive zero-shot capabilities on a thorough evaluation encompassing 14 datasets over 4 tasks, from recognizing to reasoning and classification to localization. Specifically, it outperforms the latest models such as SegEarth-OV, GeoRSCLIP, and VHM by a large margin, i.e., 11.95%, 8.04% and 3.55% on average respectively. The code is publicly available to facilitate further research at https: //github.com/zqcrafts/SkySense-O.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3cab2cff-b19f-4c58-b111-3ce1805c7b01Cited by top-tier papers8
- TESSERA: Temporal Embeddings of Surface Spectra for Earth Representation and AnalysisZhengpeng Feng, Clement Atzberger, Sadiq Jaffer, Jovana Knezevic et al.CVPR 2026 · 61 citations
- Towards Faithful Reasoning in Remote Sensing: A Perceptually-Grounded GeoSpatial Chain-of-Thought for Vision-Language ModelsJiaqi Liu, Lang Sun, Ronghao Fu, Bo YangICLR 2026 · 22 citations
- GeoMMBench and GeoMMAgent: Toward Expert-Level Multimodal Intelligence in Geoscience and Remote SensingAoran Xiao, Shihao Cheng, Yonghao Xu, Yexian Ren et al.CVPR 2026 · 6 citations
- Urban Socio-Semantic Segmentation with Vision-Language ReasoningYu Wang, Yi Wang, Rui Dai, Yujie Wang et al.ICLR 2026 · 4 citations
- SkySense-VITA: Towards Universal In-context Segmentation of Multi-modal Remote Sensing ImageryKang Wu, Lei Yu, Junwei Luo, Bo Dang et al.CVPR 2026 · 1 citation
Builds on37
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Scale-MAE: A Scale-Aware Masked Autoencoder for Multiscale Geospatial Representation LearningColorado J. Reed, Ritwik Gupta, Shufan Li, Sarah Brockman et al.ICCV 2023 · 373 citations
- SkyScript: A Large and Semantically Diverse Vision-Language Dataset for Remote SensingZhecheng Wang, Rajanie Prabha, Tianyuan Huang, Jiajun Wu et al.AAAI 2024 · 167 citations
Related papers
- A Simple Framework for Open-Vocabulary Zero-Shot SegmentationThomas Stegmüller, Tim Lebailly, Nikola Dukic, Behzad Bozorgtabar et al.ICLR 2025
- Learning to Generate Text-Grounded Mask for Open-World Semantic Segmentation from Only Image-Text PairsJunbum Cha, Jonghwan Mun, Byungseok RohCVPR 2023
- CityVG: Contrastive Fine-Tuning and Reward-Based Chain-of-Thought Reasoning for Zero-Shot City-Scale 3D Visual GroundingJianjun Zhang, Hanli WangACL 2026
- Recognizing Unseen Objects via Multimodal Intensive Knowledge Graph PropagationLikang Wu, Zhi Li, Hongke Zhao, Zhefeng Wang et al.KDD 2023 · 4 citations
- ZeroCap: Zero-Shot Image-to-Text Generation for Visual-Semantic ArithmeticYoad Tewel, Yoav Shalev, Idan Schwartz, Lior WolfCVPR 2022 · 129 citations
