ORSATR-X: A Foundation Model based on Differential-and-Excitation Networks for Optical Remote Sensing Object Recognition
Canyu Mo, Yongxiang Liu, Jiehua Zhang, Zilong Yu, Zhen Liu, Tianpeng Liu, Li Liu
Abstract
Recent advances in Remote Sensing Foundation Models (RSFMs) have demonstrated considerable potential for Earth Observation (EO) tasks. While adopting natural image foundation models (e.g., DINO) provides a dataefficient strategy for building RSFMs, their strong generalization capability does not fully transfer to complex remote sensing (RS) scenarios due to severe background interference, notably in perceiving challenging targets like low-contrast objects. To this end, we propose ORSATR-X, a novel RSFM that effectively integrates the generalizable representations of DINOv3 with a dedicated mechanism for exciting local contrast information. ORSATR-X comprises two core components: (1) a DINOv3 encoder, which provides rich feature representation under limited RS pre-training data, and (2) a carefully designed side network incorporating a Weber Local Adapter (WLA) and a Multiscale Aggregation Module (MSAM). The WLA enhances discriminability of low-contrast boundaries in complex scenes through center-surround contrast and directional gradient information enhancement, while the MSAM handles inherent object scale variations in RS imagery by adaptive aggregation of features across multiple scales. Furthermore, we pre-train the side network using an efficient self-supervised distillation strategy. Extensive experiments on scene classification, object detection, and semantic segmentation demonstrate that ORSATR-X achieves state-ofthe-art performance among existing RSFMs, demonstrating the effectiveness of our design.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b476b69d-c44f-481c-bf73-289134e6441aBuilds on24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- SimMIM: a Simple Framework for Masked Image ModelingZhenda Xie, Zheng Zhang, Yue Cao, Yutong Lin et al.CVPR 2022 · 1,129 citations
- SatMAE: Pre-training Transformers for Temporal and Multi-Spectral Satellite ImageryYezhen Cong, Samar Khanna, Chenlin Meng, Patrick Liu et al.NeurIPS 2022 · 707 citations
Related papers
- Small but Mighty: Dynamic Wavelet Expert-Guided Fine-Tuning of Large-Scale Models for Optical Remote Sensing Object SegmentationYanguang Sun, Chao Wang, Jian Yang, Lei LuoAAAI 2026 · 2 citations
- Brewing Stronger Features: Dual-Teacher Distillation for Multispectral Earth ObservationFilip Wolf, Blaz Rolih, Luka Cehovin ZajcCVPR 2026 · 4 citations
- ROS-SAM: High-Quality Interactive Segmentation for Remote Sensing Moving ObjectZhe Shan, Yang Liu, Lei Zhou, Cheng Yan et al.CVPR 2025
- SkySense V2: A Unified Foundation Model for Multi-Modal Remote SensingYingying Zhang, Lixiang Ru, Kang Wu, Lei Yu et al.ICCV 2025 · 12 citations
- DINO Eats CLIP: Adapting Beyond Knowns for Open-set 3D Object RetrievalXinwei He, Yansong Zheng, Qianru Han, Zhichuan Wang et al.CVPR 2026
