Cross-Scale MAE: A Tale of Multiscale Exploitation in Remote Sensing
Maofeng Tang, Andrei Cozma, Konstantinos Georgiou, Hairong Qi
Abstract
Remote sensing images present unique challenges to image analysis due to the extensive geographic coverage, hardware limitations, and misaligned multi-scale images. This paper revisits the classical multi-scale representation learning problem but under the general framework of self-supervised learning for remote sensing image understanding. We present Cross-Scale MAE, a self-supervised model built upon the Masked Auto-Encoder (MAE). During pre-training, Cross-Scale MAE employs scale augmentation techniques and enforces cross-scale consistency constraints through both contrastive and generative losses to ensure consistent and meaningful representations well-suited for a wide range of downstream tasks. Further, our implementation leverages the xFormers library to accelerate network pre-training on a single GPU while maintaining the quality of learned representations. Experimental evaluations demonstrate that Cross-Scale MAE exhibits superior performance compared to standard MAE and other state-of-the-art remote sensing MAE methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a2c78da1-3698-435d-b7d6-81bccf73dc6cCited by top-tier papers15
- TESSERA: Temporal Embeddings of Surface Spectra for Earth Representation and AnalysisZhengpeng Feng, Clement Atzberger, Sadiq Jaffer, Jovana Knezevic et al.CVPR 2026 · 61 citations
- TerraFM: A Scalable Foundation Model for Unified Multisensor Earth ObservationMuhammad Sohail Danish, Muhammad Akhtar Munir, Syed Roshaan Ali Shah, Muhammad Haris Khan et al.ICLR 2026 · 30 citations
- RoMA: Scaling up Mamba-based Foundation Models for Remote SensingFengxiang Wang, Yulin Wang, Mingshuo Chen, Haotian Wang et al.NeurIPS 2025 · 14 citations
- GeoLink: Empowering Remote Sensing Foundation Model with OpenStreetMap DataLubin Bai, Xiuyuan Zhang, Siqi Zhang, Zepeng Zhang et al.NeurIPS 2025 · 10 citations
- Towards a Unified Copernicus Foundation Model for Earth VisionYi Wang, Zhitong Xiong, Chenying Liu, Adam J. Stewart et al.ICCV 2025 · 7 citations
Builds on14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Barlow Twins: Self-Supervised Learning via Redundancy ReductionJure Zbontar, Li Jing, Ishan Misra, Yann LeCun et al.ICML 2021 · 2,942 citations
Related papers
- Rethinking Transformers Pre-training for Multi-Spectral Satellite ImageryMubashir Noman, Muzammal Naseer, Hisham Cholakkal, Rao Muhammad Anwer et al.CVPR 2024 · 51 citations
- S2MAE: A Spatial-Spectral Pretraining Foundation Model for Spectral Remote Sensing DataXuyang Li, Danfeng Hong, Jocelyn ChanussotCVPR 2024
- Scale-MAE: A Scale-Aware Masked Autoencoder for Multiscale Geospatial Representation LearningColorado J. Reed, Ritwik Gupta, Shufan Li, Sarah Brockman et al.ICCV 2023 · 373 citations
- NeighborMAE: Exploiting Spatial Dependencies between Neighboring Earth Observation Images in Masked Autoencoders PretrainingLiang Zeng, Valerio Marsocci, Wufan Zhao, Andrea Nascetti et al.CVPR 2026 · 1 citation
- Self-Supervised Representation Learning from Arbitrary ScenariosZhaowen Li, Yousong Zhu, Zhiyang Chen, Zongxin Gao et al.CVPR 2024
