SkySense V2: A Unified Foundation Model for Multi-Modal Remote Sensing
Yingying Zhang, Lixiang Ru, Kang Wu, Lei Yu, Lei Liang, Yansheng Li, Jingdong Chen
Abstract
The multi-modal remote sensing foundation model (MM-RSFM) has significantly advanced various Earth observation tasks, such as urban planning, environmental monitoring, and natural disaster management. However, most existing approaches generally require the training of separate backbone networks for each data modality, leading to redundancy and inefficient parameter utilization. Moreover, prevalent pre-training methods typically apply self-supervised learning (SSL) techniques from natural images without adequately accommodating the characteristics of remote sensing (RS) images, such as the complicated semantic distribution within a single RS image. In this work, we present SkySense V2, a unified MM-RSFM that employs a single transformer backbone to handle multiple modalities. This backbone is pre-trained with a novel SSL strategy tailored to the distinct traits of RS data. In particular, SkySense V2 incorporates an innovative adaptive patch merging module and learnable modality prompt tokens to address challenges related to varying resolutions and limited feature diversity across modalities. In additional, we incorporate the mixture of experts (MoE) module to further enhance the performance of the foundation model. SkySense V2 demonstrates impressive generalization abilities through an extensive evaluation involving 16 datasets over 7 tasks, outperforming SkySense by an average of 1.8 points.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 302a99bf-277a-43dd-9459-7976b2aab0e9Cited by top-tier papers8
- TESSERA: Temporal Embeddings of Surface Spectra for Earth Representation and AnalysisZhengpeng Feng, Clement Atzberger, Sadiq Jaffer, Jovana Knezevic et al.CVPR 2026 · 61 citations
- Brewing Stronger Features: Dual-Teacher Distillation for Multispectral Earth ObservationFilip Wolf, Blaz Rolih, Luka Cehovin ZajcCVPR 2026 · 4 citations
- MaRS: A Multi-modality Very-high-resolution Remote Sensing Foundation Model with Cross-Granularity Meta-Modality LearningRuoyu Yang, Yinhe Liu, Heng Yan, Yiheng Zhou et al.AAAI 2026 · 3 citations
- Local Precise Refinement: A Dual-Gated Mixture-of-Experts for Enhancing Foundation Model Generalization against Spectral ShiftsXi Chen, Maojun Zhang, Yu Liu, Shen YanCVPR 2026 · 3 citations
- ORSATR-X: A Foundation Model based on Differential-and-Excitation Networks for Optical Remote Sensing Object RecognitionCanyu Mo, Yongxiang Liu, Jiehua Zhang, Zilong Yu et al.CVPR 2026 · 1 citation
Builds on31
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
Related papers
- Any-Optical-Model: A Universal Foundation Model for Optical Remote SensingXuyang Li, Chenyu Li, Danfeng HongAAAI 2026
- SkySense: A Multi-Modal Remote Sensing Foundation Model Towards Universal Interpretation for Earth Observation ImageryXin Guo, Jiangwei Lao, Bo Dang, Yingying Zhang et al.CVPR 2024
- EVE: Efficient Vision-Language Pre-training with Masked Prediction and Modality-Aware MoEJunyi Chen, Longteng Guo, Jia Sun, Shuai Shao et al.AAAI 2024 · 25 citations
- SkySense-VITA: Towards Universal In-context Segmentation of Multi-modal Remote Sensing ImageryKang Wu, Lei Yu, Junwei Luo, Bo Dang et al.CVPR 2026 · 1 citation
- S5: Scalable Semi-Supervised Semantic Segmentation in Remote SensingLiang Lv, Di Wang, Jing Zhang, Lefei ZhangAAAI 2026 · 4 citations
