Cross-Scale MAE: A Tale of Multiscale Exploitation in Remote Sensing
Maofeng Tang, Andrei Cozma, Konstantinos Georgiou, Hairong Qi
摘要
Remote sensing images present unique challenges to image analysis due to the extensive geographic coverage, hardware limitations, and misaligned multi-scale images. This paper revisits the classical multi-scale representation learning problem but under the general framework of self-supervised learning for remote sensing image understanding. We present Cross-Scale MAE, a self-supervised model built upon the Masked Auto-Encoder (MAE). During pre-training, Cross-Scale MAE employs scale augmentation techniques and enforces cross-scale consistency constraints through both contrastive and generative losses to ensure consistent and meaningful representations well-suited for a wide range of downstream tasks. Further, our implementation leverages the xFormers library to accelerate network pre-training on a single GPU while maintaining the quality of learned representations. Experimental evaluations demonstrate that Cross-Scale MAE exhibits superior performance compared to standard MAE and other state-of-the-art remote sensing MAE methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- TESSERA: Temporal Embeddings of Surface Spectra for Earth Representation and AnalysisZhengpeng Feng, Clement Atzberger, Sadiq Jaffer, Jovana Knezevic 等CVPR 2026 · 被引用 61 次
- TerraFM: A Scalable Foundation Model for Unified Multisensor Earth ObservationMuhammad Sohail Danish, Muhammad Akhtar Munir, Syed Roshaan Ali Shah, Muhammad Haris Khan 等ICLR 2026 · 被引用 30 次
- RoMA: Scaling up Mamba-based Foundation Models for Remote SensingFengxiang Wang, Yulin Wang, Mingshuo Chen, Haotian Wang 等NeurIPS 2025 · 被引用 14 次
- GeoLink: Empowering Remote Sensing Foundation Model with OpenStreetMap DataLubin Bai, Xiuyuan Zhang, Siqi Zhang, Zepeng Zhang 等NeurIPS 2025 · 被引用 10 次
- Towards a Unified Copernicus Foundation Model for Earth VisionYi Wang, Zhitong Xiong, Chenying Liu, Adam J. Stewart 等ICCV 2025 · 被引用 7 次
它引用的顶会 Paper14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Barlow Twins: Self-Supervised Learning via Redundancy ReductionJure Zbontar, Li Jing, Ishan Misra, Yann LeCun 等ICML 2021 · 被引用 2,942 次
相关 Paper
- Rethinking Transformers Pre-training for Multi-Spectral Satellite ImageryMubashir Noman, Muzammal Naseer, Hisham Cholakkal, Rao Muhammad Anwer 等CVPR 2024 · 被引用 51 次
- S2MAE: A Spatial-Spectral Pretraining Foundation Model for Spectral Remote Sensing DataXuyang Li, Danfeng Hong, Jocelyn ChanussotCVPR 2024
- Scale-MAE: A Scale-Aware Masked Autoencoder for Multiscale Geospatial Representation LearningColorado J. Reed, Ritwik Gupta, Shufan Li, Sarah Brockman 等ICCV 2023 · 被引用 373 次
- NeighborMAE: Exploiting Spatial Dependencies between Neighboring Earth Observation Images in Masked Autoencoders PretrainingLiang Zeng, Valerio Marsocci, Wufan Zhao, Andrea Nascetti 等CVPR 2026 · 被引用 1 次
- Self-Supervised Representation Learning from Arbitrary ScenariosZhaowen Li, Yousong Zhu, Zhiyang Chen, Zongxin Gao 等CVPR 2024
