CROMA: Remote Sensing Representations with Contrastive Radar-Optical Masked Autoencoders
Anthony Fuller, Koreen Millard, James R. Green
Abstract
A vital and rapidly growing application, remote sensing offers vast yet sparsely labeled, spatially aligned multimodal data; this makes self-supervised learning algorithms invaluable. We present CROMA: a framework that combines contrastive and reconstruction self-supervised objectives to learn rich unimodal and multimodal representations. Our method separately encodes masked-out multispectral optical and synthetic aperture radar samples -- aligned in space and time -- and performs cross-modal contrastive learning. Another encoder fuses these sensors, producing joint multimodal encodings that are used to predict the masked patches via a lightweight decoder. We show that these objectives are complementary when leveraged on spatially aligned multimodal data. We also introduce X- and 2D-ALiBi, which spatially biases our cross- and self-attention matrices. These strategies improve representations and allow our models to effectively extrapolate to images up to 17.6x larger at test-time. CROMA outperforms the current SoTA multispectral model, evaluated on: four classification benchmarks -- finetuning (avg. 1.8%), linear (avg. 2.4%) and nonlinear (avg. 1.4%) probing, kNN classification (avg. 3.5%), and K-means clustering (avg. 8.4%); and three segmentation benchmarks (avg. 6.4%). CROMA's rich, optionally multimodal representations can be widely leveraged across remote sensing applications.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f165f044-902e-4136-9f22-cc02a6720a03Cited by top-tier papers31
- TESSERA: Temporal Embeddings of Surface Spectra for Earth Representation and AnalysisZhengpeng Feng, Clement Atzberger, Sadiq Jaffer, Jovana Knezevic et al.CVPR 2026 · 61 citations
- TerraMind: Large-Scale Generative Multimodality for Earth ObservationJohannes Jakubik, Felix Yang, Benedikt Blumenstiel, Erik Scheurer et al.ICCV 2025 · 43 citations
- TerraFM: A Scalable Foundation Model for Unified Multisensor Earth ObservationMuhammad Sohail Danish, Muhammad Akhtar Munir, Syed Roshaan Ali Shah, Muhammad Haris Khan et al.ICLR 2026 · 30 citations
- OlmoEarth: Stable Latent Image Modeling for Multimodal Earth ObservationHenry Herzog, Favyen Bastani, Yawen Zhang, Gabriel Tseng et al.CVPR 2026 · 28 citations
- RoMA: Scaling up Mamba-based Foundation Models for Remote SensingFengxiang Wang, Yulin Wang, Mingshuo Chen, Haotian Wang et al.NeurIPS 2025 · 14 citations
Builds on42
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
Related papers
- Galileo: Learning Global & Local Features of Many Remote Sensing ModalitiesGabriel Tseng, Anthony Fuller, Marlena Reil, Henry Herzog et al.ICML 2025
- CoMAE: Single Model Hybrid Pre-training on Small-Scale RGB-D DatasetsJiange Yang, Sheng Guo, Gangshan Wu, Limin WangAAAI 2023 · 12 citations
- Cross-Scale MAE: A Tale of Multiscale Exploitation in Remote SensingMaofeng Tang, Andrei Cozma, Konstantinos Georgiou, Hairong QiNeurIPS 2023 · 88 citations
- COCOA: Cross Modality Contrastive Learning for Sensor DataShohreh Deldari, Hao Xue, Aaqib Saeed, Daniel V. Smith et al.UbiComp 2022 · 88 citations
- SkySense: A Multi-Modal Remote Sensing Foundation Model Towards Universal Interpretation for Earth Observation ImageryXin Guo, Jiangwei Lao, Bo Dang, Yingying Zhang et al.CVPR 2024
