Scale-MAE: A Scale-Aware Masked Autoencoder for Multiscale Geospatial Representation Learning
Colorado J. Reed, Ritwik Gupta, Shufan Li, Sarah Brockman, Christopher Funk, Brian Clipp, Kurt Keutzer, Salvatore Candido, Matt Uyttendaele, Trevor Darrell
Abstract
Large, pretrained models are commonly finetuned with imagery that is heavily augmented to mimic different conditions and scales, with the resulting models used for various tasks with imagery from a range of spatial scales. Such models overlook scale-specific information in the data for scale-dependent domains, such as remote sensing. In this paper, we present Scale-MAE, a pretraining method that explicitly learns relationships between data at different, known scales throughout the pretraining process. Scale-MAE pretrains a network by masking an input image at a known input scale, where the area of the Earth covered by the image determines the scale of the ViT positional encoding, not the image resolution. Scale-MAE encodes the masked image with a standard ViT backbone, and then decodes the masked image through a bandpass filter to reconstruct low/high frequency images at lower/higher scales. We find that tasking the network with reconstructing both low/high frequency images leads to robust multiscale representations for remote sensing imagery. Scale-MAE achieves an average of a 2.4 -5.6% non-parametric kNN classification improvement across eight remote sensing datasets compared to current state-of-the-art and obtains a 0.9 mIoU to 1.7 mIoU improvement on the SpaceNet building segmentation transfer task for a range of evaluation scales. * Denotes co-first authorship. Co-first authors will prioritize their names on their resumes/websites.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 47ff35db-8700-4a57-8d27-65fed2089e79Cited by top-tier papers52
- ClimaX: A foundation model for weather and climateTung Nguyen, Johannes Brandstetter, Ashish Kapoor, Jayesh K. Gupta et al.ICML 2023 · 426 citations
- SkyScript: A Large and Semantically Diverse Vision-Language Dataset for Remote SensingZhecheng Wang, Rajanie Prabha, Tianyuan Huang, Jiajun Wu et al.AAAI 2024 · 167 citations
- Cross-Scale MAE: A Tale of Multiscale Exploitation in Remote SensingMaofeng Tang, Andrei Cozma, Konstantinos Georgiou, Hairong QiNeurIPS 2023 · 88 citations
- TESSERA: Temporal Embeddings of Surface Spectra for Earth Representation and AnalysisZhengpeng Feng, Clement Atzberger, Sadiq Jaffer, Jovana Knezevic et al.CVPR 2026 · 61 citations
- SwitchTab: Switched Autoencoders Are Effective Tabular LearnersJing Wu, Suiyao Chen, Qi Zhao, Renat Sergazinov et al.AAAI 2024 · 60 citations
Builds on11
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- Multiscale Vision TransformersHaoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li et al.ICCV 2021 · 1,611 citations
- Data-Efficient Image Recognition with Contrastive Predictive CodingOlivier J. HénaffICML 2020 · 1,553 citations
- SatMAE: Pre-training Transformers for Temporal and Multi-Spectral Satellite ImageryYezhen Cong, Samar Khanna, Chenlin Meng, Patrick Liu et al.NeurIPS 2022 · 707 citations
Related papers
- Rethinking Transformers Pre-training for Multi-Spectral Satellite ImageryMubashir Noman, Muzammal Naseer, Hisham Cholakkal, Rao Muhammad Anwer et al.CVPR 2024 · 51 citations
- SparseMAE: Sparse Training Meets Masked AutoencodersAojun Zhou, Yang Li, Zipeng Qin, Jianbo Liu et al.ICCV 2023 · 8 citations
- S2MAE: A Spatial-Spectral Pretraining Foundation Model for Spectral Remote Sensing DataXuyang Li, Danfeng Hong, Jocelyn ChanussotCVPR 2024
- MCMAE: Masked Convolution Meets Masked AutoencodersPeng Gao, Teli Ma, Hongsheng Li, Ziyi Lin et al.NeurIPS 2022 · 84 citations
- Masked Autoencoders Are Scalable Vision LearnersKaiming He, Xinlei Chen, Saining Xie, Yanghao Li et al.CVPR 2022
