Any-Optical-Model: A Universal Foundation Model for Optical Remote Sensing
Xuyang Li, Chenyu Li, Danfeng Hong
Abstract
Optical satellites, with their diverse band layouts and ground sampling distances, supply indispensable evidence for tasks ranging from ecosystem surveillance to emergency response. However, significant discrepancies in band composition and spatial resolution across different optical sensors present major challenges for existing Remote Sensing Foundation Models (RSFMs). These models are typically pretrained on fixed band configurations and resolutions, making them vulnerable to real world scenarios involving missing bands, cross sensor fusion, and unseen spatial scales, thereby limiting their generalization and practical deployment. To address these limitations, we propose Any Optical Model (AOM), a universal RSFM explicitly designed to accommodate arbitrary band compositions, sensor types, and resolution scales. To preserve distinctive spectral characteristics even when bands are missing or newly introduced, AOM introduces a spectrum-independent tokenizer that assigns each channel a dedicated band embedding, enabling explicit encoding of spectral identity. To effectively capture texture and contextual patterns from sub-meter to hundred-meter imagery, we design a multi-scale adaptive patch embedding mechanism that dynamically modulates the receptive field. Furthermore, to maintain global semantic consistency across varying resolutions, AOM incorporates a multi-scale semantic alignment mechanism alongside a channel-wise self-supervised masking and reconstruction pretraining strategy that jointly models spectral-spatial relationships. Extensive experiments on over 10 public datasets, including those from Sentinel-2, Landsat, and HLS, demonstrate that AOM consistently achieves state-of-the-art (SOTA) performance under challenging conditions such as band missing, cross sensor, and cross resolution settings. These results highlight AOM as a crucial step toward building truly general-purpose RSFMs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4cddf64f-0f93-48d1-958f-8d37c66b19eaBuilds on13
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- SatMAE: Pre-training Transformers for Temporal and Multi-Spectral Satellite ImageryYezhen Cong, Samar Khanna, Chenlin Meng, Patrick Liu et al.NeurIPS 2022 · 707 citations
- Scale-MAE: A Scale-Aware Masked Autoencoder for Multiscale Geospatial Representation LearningColorado J. Reed, Ritwik Gupta, Shufan Li, Sarah Brockman et al.ICCV 2023 · 373 citations
- Towards Geospatial Foundation Models via Continual PretrainingMatías Mendieta, Boran Han, Xingjian Shi, Yi Zhu et al.ICCV 2023 · 140 citations
- Rethinking Transformers Pre-training for Multi-Spectral Satellite ImageryMubashir Noman, Muzammal Naseer, Hisham Cholakkal, Rao Muhammad Anwer et al.CVPR 2024 · 51 citations
Related papers
- SkySense V2: A Unified Foundation Model for Multi-Modal Remote SensingYingying Zhang, Lixiang Ru, Kang Wu, Lei Yu et al.ICCV 2025 · 12 citations
- RobSense: A Robust Multi-modal Foundation Model for Remote Sensing with Static, Temporal, and Incomplete Data AdaptabilityMinh Kha Do, Kang Han, Phu Lai, Khoa T. Phan et al.CVPR 2025
- SMARTIES: Spectrum-Aware Multi-Sensor Auto-Encoder for Remote Sensing ImagesGencer Sumbul, Chang Xu, Emanuele Dalsasso, Devis TuiaICCV 2025 · 3 citations
- PhySwin: An Efficient and Physically-Informed Foundation Model for Multispectral Earth ObservationChong Tang, Joseph Powell, Dirk Koch, Robert Mullins et al.NeurIPS 2025 · 1 citation
- Bridging Remote Sensors with Multisensor Geospatial Foundation ModelsBoran Han, Shuai Zhang, Xingjian Shi, Markus ReichsteinCVPR 2024
