SkySense-VITA: Towards Universal In-context Segmentation of Multi-modal Remote Sensing Imagery
Kang Wu, Lei Yu, Junwei Luo, Bo Dang, Junjian Zhang, Xiangyuan Cai, Hongwei Hu, Jingdong Chen, Yansheng Li
摘要
While recent foundation models for remote sensing segmentation have shown notable progress, they still fall short in processing diverse multi-modal inputs, synergizing complementary prompt types, and leveraging semantic hierarchies. To address these limitations, we introduce SkySense-VITA, a unified in-context segmentation model, which synergistically processes both optical and Synthetic Aperture Radar (SAR) imagery using VIsual, TextuAl, or fused prompts. Based on a novel prompt-and-prediction decoupling strategy, we propose the VITA-Former and VITA-Decoder to decouple multi-modal prompt fusion and prediction process, allowing the model to flexibly support visual-only, textualonly, and fused prompt modes. We train SkySense-VITA with a progressive two-stage strategy: a first stage of Image-Level Alignment Pretraining featuring optical-SAR alignment, and a second stage of Pixel-Level In-context Pretraining using Semantic Granularity Annealing (SGA), a coarseto-fine curriculum that enables robust hierarchical learning. To support this training, we introduce our new largescale, multi-modal Sky-VT-300k dataset. Extensive experiments show SkySense-VITA establishes a new state-of-theart (SOTA) on 18 datasets, with an average performance lead of over 10% mean Intersection over Union (mIoU).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty 等NeurIPS 2021 · 被引用 2,985 次
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 被引用 2,196 次
- SatMAE: Pre-training Transformers for Temporal and Multi-Spectral Satellite ImageryYezhen Cong, Samar Khanna, Chenlin Meng, Patrick Liu 等NeurIPS 2022 · 被引用 707 次
相关 Paper
- T-APT: Text-Guided Modality-Aware Prompt Tuning for Arbitrary Multimodal Remote Sensing Data Joint ClassificationQinghao Gao, Jiahui Qu, Wenqian DongAAAI 2026
- SkySense V2: A Unified Foundation Model for Multi-Modal Remote SensingYingying Zhang, Lixiang Ru, Kang Wu, Lei Yu 等ICCV 2025 · 被引用 12 次
- Unified Open-World Segmentation with Multi-Modal PromptsYang Liu, Yufei Yin, Chenchen Jing, Muzhi Zhu 等ICCV 2025 · 被引用 8 次
- MM-OVSeg: Multimodal Optical-SAR Fusion for Open-Vocabulary Segmentation in Remote SensingYimin Wei, Aoran Xiao, Hongruixuan Chen, Junshi Xia 等CVPR 2026 · 被引用 6 次
- RS2-SAM2: Customized SAM2 for Referring Remote Sensing Image SegmentationFu Rong, Meng Lan, Qian Zhang, Lefei ZhangAAAI 2026 · 被引用 1 次
