SCORE: Scene Context Matters in Open-Vocabulary Remote Sensing Instance Segmentation
Shiqi Huang, Shuting He, Huaiyuan Qin, Bihan Wen
Abstract
Most existing remote sensing instance segmentation approaches are designed for close-vocabulary prediction, limiting their ability to recognize novel categories or generalize across datasets. This restricts their applicability in diverse Earth observation scenarios. To address this, we introduce open-vocabulary (OV) learning for remote sensing instance segmentation. While current OV segmentation models perform well on natural image datasets, their direct application to remote sensing faces challenges such as diverse landscapes, seasonal variations, and the presence of small or ambiguous objects in aerial imagery. To overcome these challenges, we propose SCORE (Scene Context matters in Open-vocabulary REmote sensing instance segmentation), a framework that integrates multigranularity scene context, i.e., regional context and global context, to enhance both visual and textual representations. Specifically, we introduce Region-Aware Integration, which refines class embeddings with regional context to improve object distinguishability. Additionally, we propose Global Context Adaptation, which enriches naive text embeddings with remote sensing global context, creating a more adaptable and expressive linguistic latent space for the classifier. We establish new benchmarks for OV remote sensing instance segmentation across diverse datasets. Experimental results demonstrate that, our proposed method achieves SOTA performance, which provides a robust solution for largescale, real-world geospatial analysis. Our code is available at https://github.com/HuangShiqi128/SCORE.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c539eaaa-1c94-44bd-9b21-3cd250210cbbCited by top-tier papers2
- TerraScope: Pixel-Grounded Visual Reasoning for Earth ObservationYan Shu, Bin Ren, Zhitong Xiong, Xiao Xiang Zhu et al.CVPR 2026 · 9 citations
- HySeg: Learning Generative Priors for Structure-Aware Remote Sensing SegmentationJie Qiu, XIN LI, Fan Yang, Yan Wang et al.CVPR 2026
Builds on23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Open-vocabulary Object Detection via Vision and Language Knowledge DistillationXiuye Gu, Tsung-Yi Lin, Weicheng Kuo, Yin CuiICLR 2022 · 1,274 citations
- Language-driven Semantic SegmentationBoyi Li, Kilian Q. Weinberger, Serge J. Belongie, Vladlen Koltun et al.ICLR 2022 · 885 citations
Related papers
- ZoRI: Towards Discriminative Zero-Shot Remote Sensing Instance SegmentationShiqi Huang, Shuting He, Bihan WenAAAI 2025
- SOAR: Semi-Supervised Open-Vocabulary Aerial Object Detection via Dual-Aware Enhanced Prior DenoisingXu Liu, Yihong Huang, Dan Zhang, Lingling Li et al.AAAI 2026
- OV3D-CG: Open-Vocabulary 3D Instance Segmentation with Contextual GuidanceMingquan Zhou, Chen He, Ruiping Wang, Xilin ChenICCV 2025 · 1 citation
- OpenRSD: Towards Open-Prompts for Object Detection in Remote Sensing ImagesZiyue Huang, Yongchao Feng, Ziqi Liu, Shuai Yang et al.ICCV 2025 · 3 citations
- Test-Time Multi-Prompt Adaptation for Open-Vocabulary Remote Sensing Image SegmentationTing Yang, Qilong Wang, Qibin Hou, Qinghua HuCVPR 2026
