Scale-MAE: A Scale-Aware Masked Autoencoder for Multiscale Geospatial Representation Learning
Colorado J. Reed, Ritwik Gupta, Shufan Li, Sarah Brockman, Christopher Funk, Brian Clipp, Kurt Keutzer, Salvatore Candido, Matt Uyttendaele, Trevor Darrell
摘要
Large, pretrained models are commonly finetuned with imagery that is heavily augmented to mimic different conditions and scales, with the resulting models used for various tasks with imagery from a range of spatial scales. Such models overlook scale-specific information in the data for scale-dependent domains, such as remote sensing. In this paper, we present Scale-MAE, a pretraining method that explicitly learns relationships between data at different, known scales throughout the pretraining process. Scale-MAE pretrains a network by masking an input image at a known input scale, where the area of the Earth covered by the image determines the scale of the ViT positional encoding, not the image resolution. Scale-MAE encodes the masked image with a standard ViT backbone, and then decodes the masked image through a bandpass filter to reconstruct low/high frequency images at lower/higher scales. We find that tasking the network with reconstructing both low/high frequency images leads to robust multiscale representations for remote sensing imagery. Scale-MAE achieves an average of a 2.4 -5.6% non-parametric kNN classification improvement across eight remote sensing datasets compared to current state-of-the-art and obtains a 0.9 mIoU to 1.7 mIoU improvement on the SpaceNet building segmentation transfer task for a range of evaluation scales. * Denotes co-first authorship. Co-first authors will prioritize their names on their resumes/websites.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper52
- ClimaX: A foundation model for weather and climateTung Nguyen, Johannes Brandstetter, Ashish Kapoor, Jayesh K. Gupta 等ICML 2023 · 被引用 426 次
- SkyScript: A Large and Semantically Diverse Vision-Language Dataset for Remote SensingZhecheng Wang, Rajanie Prabha, Tianyuan Huang, Jiajun Wu 等AAAI 2024 · 被引用 167 次
- Cross-Scale MAE: A Tale of Multiscale Exploitation in Remote SensingMaofeng Tang, Andrei Cozma, Konstantinos Georgiou, Hairong QiNeurIPS 2023 · 被引用 88 次
- TESSERA: Temporal Embeddings of Surface Spectra for Earth Representation and AnalysisZhengpeng Feng, Clement Atzberger, Sadiq Jaffer, Jovana Knezevic 等CVPR 2026 · 被引用 61 次
- SwitchTab: Switched Autoencoders Are Effective Tabular LearnersJing Wu, Suiyao Chen, Qi Zhao, Renat Sergazinov 等AAAI 2024 · 被引用 60 次
它引用的顶会 Paper11
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer 等CVPR 2022 · 被引用 6,782 次
- Multiscale Vision TransformersHaoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li 等ICCV 2021 · 被引用 1,611 次
- Data-Efficient Image Recognition with Contrastive Predictive CodingOlivier J. HénaffICML 2020 · 被引用 1,553 次
- SatMAE: Pre-training Transformers for Temporal and Multi-Spectral Satellite ImageryYezhen Cong, Samar Khanna, Chenlin Meng, Patrick Liu 等NeurIPS 2022 · 被引用 707 次
相关 Paper
- Rethinking Transformers Pre-training for Multi-Spectral Satellite ImageryMubashir Noman, Muzammal Naseer, Hisham Cholakkal, Rao Muhammad Anwer 等CVPR 2024 · 被引用 51 次
- SparseMAE: Sparse Training Meets Masked AutoencodersAojun Zhou, Yang Li, Zipeng Qin, Jianbo Liu 等ICCV 2023 · 被引用 8 次
- S2MAE: A Spatial-Spectral Pretraining Foundation Model for Spectral Remote Sensing DataXuyang Li, Danfeng Hong, Jocelyn ChanussotCVPR 2024
- MCMAE: Masked Convolution Meets Masked AutoencodersPeng Gao, Teli Ma, Hongsheng Li, Ziyi Lin 等NeurIPS 2022 · 被引用 84 次
- Masked Autoencoders Are Scalable Vision LearnersKaiming He, Xinlei Chen, Saining Xie, Yanghao Li 等CVPR 2022
