SatCLIP: Global, General-Purpose Location Embeddings with Satellite Imagery
Konstantin Klemmer, Esther Rolf, Caleb Robinson, Lester Mackey, Marc Rußwurm
Abstract
Geographic information is essential for modeling tasks in fields ranging from ecology to epidemiology. However, extracting relevant location characteristics for a given task can be challenging, often requiring expensive data fusion or distillation from massive global imagery datasets. To address this challenge, we introduce Satellite Contrastive Location-Image Pretraining (SatCLIP). This global, general-purpose geographic location encoder learns an implicit representation of locations by matching CNN and ViT inferred visual patterns of openly available satellite imagery with their geographic coordinates. The resulting SatCLIP location encoder efficiently summarizes the characteristics of any given location for convenient use in downstream tasks. In our experiments, we use SatCLIP embeddings to improve prediction performance on nine diverse location-dependent tasks including temperature prediction, animal recognition, and population density estimation. Across tasks, SatCLIP consistently outperforms alternative location encoders and improves geographic generalization by encoding visual similarities of spatially distant environments. These results demonstrate the potential of vision-location models to learn meaningful representations of our planet from the vast, varied, and largely untapped modalities of geospatial data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d4fbd998-420f-4b4d-93b9-e4ea7f67481bCited by top-tier papers26
- TESSERA: Temporal Embeddings of Surface Spectra for Earth Representation and AnalysisZhengpeng Feng, Clement Atzberger, Sadiq Jaffer, Jovana Knezevic et al.CVPR 2026 · 61 citations
- Combining Observational Data and Language for Species Range EstimationMax Hamilton, Christian Lange, Elijah Cole, Alexander Shepard et al.NeurIPS 2024 · 18 citations
- Measuring the Intrinsic Dimension of Earth RepresentationsArjun Rao, Marc Rußwurm, Konstantin Klemmer, Esther RolfICLR 2026 · 13 citations
- Nature Makes No Leaps: Building Continuous Location Embeddings with Satellite Imagery from the WebXixuan Hao, Wei Chen, Xingchen Zou, Yuxuan LiangWWW 2025 · 11 citations
- Towards a Unified Copernicus Foundation Model for Earth VisionYi Wang, Zhitong Xiong, Chenying Liu, Adam J. Stewart et al.ICCV 2025 · 7 citations
Builds on12
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Implicit Neural Representations with Periodic Activation FunctionsVincent Sitzmann, Julien N. P. Martel, Alexander W. Bergman, David B. Lindell et al.NeurIPS 2020 · 4,008 citations
- Presence-Only Geographical Priors for Fine-Grained Image ClassificationOisin Mac Aodha, Elijah Cole, Pietro PeronaICCV 2019 · 206 citations
- Multi-Scale Representation Learning for Spatial Feature Distributions using Grid CellsGengchen Mai, Krzysztof Janowicz, Bo Yan, Rui Zhu et al.ICLR 2020 · 161 citations
Related papers
- RANGE: Retrieval Augmented Neural Fields for Multi-Resolution Geo-EmbeddingsAayush Dhakal, Srikumar Sastry, Subash Khanal, Adeel Ahmad et al.CVPR 2025
- GeoCLIP: Clip-Inspired Alignment between Locations and Images for Effective Worldwide Geo-localizationVicente Vivanco Cepeda, Gaurav Kumar Nayak, Mubarak ShahNeurIPS 2023 · 303 citations
- Beyond What's Shared: Recovering Lost Unique Information from Intermediate Layers to Boost Multimodal Geo-Foundation ModelsJangHyeon Lee, Philipe Ambrozio Dias, Yao-Yi Chiang, Dalton LungaCVPR 2026
- How Much Can CLIP Benefit Vision-and-Language Tasks?Sheng Shen, Liunian Harold Li, Hao Tan, Mohit Bansal et al.ICLR 2022 · 503 citations
- Contrastive Localized Language-Image Pre-TrainingHong-You Chen, Zhengfeng Lai, Haotian Zhang, Xinze Wang et al.ICML 2025
