SMARTIES: Spectrum-Aware Multi-Sensor Auto-Encoder for Remote Sensing Images
Gencer Sumbul, Chang Xu, Emanuele Dalsasso, Devis Tuia
Abstract
From optical sensors to microwave radars, leveraging the complementary strengths of remote sensing (RS) sensors is crucial for achieving dense spatio-temporal monitoring of our planet. In contrast, recent deep learning models, whether task-specific or foundational, are often specific to single sensors or to fixed combinations: adapting such models to different sensory inputs requires both architectural changes and re-training, limiting scalability and generalization across multiple RS sensors. On the contrary, a single model able to modulate its feature representations to accept diverse sensors as input would pave the way to agile and flexible multi-sensor RS data processing. To address this, we introduce SMARTIES, a generic and versatile foundation model lifting sensor-specific/dependent efforts and enabling scalability and generalization to diverse RS sensors: SMARTIES projects data from heterogeneous sensors into a shared spectrum-aware space, enabling the use of arbitrary combinations of bands both for training and inference. To obtain sensor-agnostic representations, we train a single, unified transformer model reconstructing masked multisensor data with cross-sensor token mixup. On both singleand multi-modal tasks across diverse sensors, SMARTIES outperforms previous models that rely on sensor-specific pretraining. Our code and pretrained models are available at https://gsumbul.qithub.io/SMARTIES.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fd3c07f7-0ded-48b8-825b-17286e603202Cited by top-tier papers2
- CARL: Camera-Agnostic Representation Learning for Spectral Image AnalysisAlexander Baumann, Leonardo Ayala, Silvia Seidlitz, Jan Sellner et al.ICLR 2026 · 2 citations
- MIAM: Modality Imbalance-Aware Masking for Multimodal Ecological ApplicationsRobin Zbinden, Wesley Monteith-Finas, Gencer Sumbul, Nina Van Tiel et al.ICLR 2026
Builds on20
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- SatMAE: Pre-training Transformers for Temporal and Multi-Spectral Satellite ImageryYezhen Cong, Samar Khanna, Chenlin Meng, Patrick Liu et al.NeurIPS 2022 · 707 citations
- Scale-MAE: A Scale-Aware Masked Autoencoder for Multiscale Geospatial Representation LearningColorado J. Reed, Ritwik Gupta, Shufan Li, Sarah Brockman et al.ICCV 2023 · 373 citations
- Seasonal Contrast: Unsupervised Pre-Training from Uncurated Remote Sensing DataOscar Mañas, Alexandre Lacoste, Xavier Giró-i-Nieto, David Vázquez et al.ICCV 2021 · 361 citations
- Geography-Aware Self-Supervised LearningKumar Ayush, Burak Uzkent, Chenlin Meng, Kumar Tanmay et al.ICCV 2021 · 304 citations
Related papers
- Any-Optical-Model: A Universal Foundation Model for Optical Remote SensingXuyang Li, Chenyu Li, Danfeng HongAAAI 2026
- Galileo: Learning Global & Local Features of Many Remote Sensing ModalitiesGabriel Tseng, Anthony Fuller, Marlena Reil, Henry Herzog et al.ICML 2025
- Hyperspectral Image Fusion with Spectral-Band and Fusion-Scale AgnosticismYujie Liang, ZiHan Cao, Liang-Jian Deng, Yang Yang et al.ICML 2026
- RAMEN: Resolution-Adjustable Multimodal Encoder for Earth ObservationNicolas Houdré, Diego Marcos, Hugo Riffaud de Turckheim, Dino Ienco et al.CVPR 2026 · 4 citations
- SkySense V2: A Unified Foundation Model for Multi-Modal Remote SensingYingying Zhang, Lixiang Ru, Kang Wu, Lei Yu et al.ICCV 2025 · 12 citations
