AnySat: One Earth Observation Model for Many Resolutions, Scales, and Modalities
Guillaume Astruc, Nicolas Gonthier, Clément Mallet, Loïc Landrieu
Abstract
Geospatial models must adapt to the diversity of Earth observation data in terms of resolutions, scales, and modalities. However, existing approaches expect fixed input configurations, which limits their practical applicability. We propose AnySat, a multimodal model based on joint embedding predictive architecture (JEPA) and scale-adaptive spatial encoders, allowing us to train a single model on highly heterogeneous data in a self-supervised manner. To demonstrate the advantages of this unified approach, we compile GeoPlex, a collection of 5 multimodal datasets with varying characteristics and 11 distinct sensors. We then train a single powerful model on these diverse datasets simultaneously. Once finetuned or probed, we reach state-of-the-art results on the test sets of GeoPlex and for 6 external datasets across various environment monitoring tasks: land cover mapping, tree species identification, crop type classification, change detection, climate type classification, and segmentation of flood, burn scar, and deforestation. Our code and models are available at https://github.com/gastruc/AnySat.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers16
- TESSERA: Temporal Embeddings of Surface Spectra for Earth Representation and AnalysisZhengpeng Feng, Clement Atzberger, Sadiq Jaffer, Jovana Knezevic et al.CVPR 2026 · 61 citations
- TerraMind: Large-Scale Generative Multimodality for Earth ObservationJohannes Jakubik, Felix Yang, Benedikt Blumenstiel, Erik Scheurer et al.ICCV 2025 · 43 citations
- OlmoEarth: Stable Latent Image Modeling for Multimodal Earth ObservationHenry Herzog, Favyen Bastani, Yawen Zhang, Gabriel Tseng et al.CVPR 2026 · 28 citations
- RoMA: Scaling up Mamba-based Foundation Models for Remote SensingFengxiang Wang, Yulin Wang, Mingshuo Chen, Haotian Wang et al.NeurIPS 2025 · 14 citations
- SkySense V2: A Unified Foundation Model for Multi-Modal Remote SensingYingying Zhang, Lixiang Ru, Kang Wu, Lei Yu et al.ICCV 2025 · 12 citations
Builds on21
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun et al.ICCV 2021 · 2,947 citations
- data2vec: A General Framework for Self-supervised Learning in Speech, Vision and LanguageAlexei Baevski, Wei-Ning Hsu, Qiantong Xu, Arun Babu et al.ICML 2022 · 1,123 citations
Related papers
- M3-JEPA: Multimodal Alignment via Multi-gate MoE based on the Joint-Embedding Predictive ArchitectureHongyang Lei, Xiaolong Cheng, Qi Qin, Dan Wang et al.ICML 2025
- Any-Optical-Model: A Universal Foundation Model for Optical Remote SensingXuyang Li, Chenyu Li, Danfeng HongAAAI 2026
- Galileo: Learning Global & Local Features of Many Remote Sensing ModalitiesGabriel Tseng, Anthony Fuller, Marlena Reil, Henry Herzog et al.ICML 2025
- DUNIA: Pixel-Sized Embeddings via Cross-Modal Alignment for Earth Observation ApplicationsIbrahim Fayad, Max Zimmer, Martin Schwartz, Fabian Gieseke et al.ICML 2025
- RAMEN: Resolution-Adjustable Multimodal Encoder for Earth ObservationNicolas Houdré, Diego Marcos, Hugo Riffaud de Turckheim, Dino Ienco et al.CVPR 2026 · 4 citations
