WildSAT: Learning Satellite Image Representations from Wildlife Observations
Rangel Daroya, Elijah Cole, Oisin Mac Aodha, Grant Van Horn, Subhransu Maji
Abstract
Species distributions encode valuable ecological and environmental information, yet their potential for guiding representation learning in remote sensing remains underexplored. We introduce WildSAT, which pairs satellite images with millions of geo-tagged wildlife observations readily-available on citizen science platforms. WildSAT employs a contrastive learning approach that jointly leverages satellite images, species occurrence maps, and textual habitat descriptions to train or fine-tune models. This approach significantly improves performance on diverse satellite image recognition tasks, outperforming both ImageNet-pretrained models and satellite-specific baselines. Additionally, by aligning visual and textual information, WildSAT enables zero-shot retrieval, allowing users to search geographic locations based on textual descriptions. WildSAT surpasses recent cross-modal learning methods, including approaches that align satellite images with ground imagery or wildlife photos, demonstrating the advantages of our approach. Finally, we analyze the impact of key design choices and highlight the broad applicability of WildSAT to remote sensing and biodiversity monitoring.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0ee6a0da-a755-4076-b785-199c6bef0e87Cited by top-tier papers4
- TESSERA: Temporal Embeddings of Surface Spectra for Earth Representation and AnalysisZhengpeng Feng, Clement Atzberger, Sadiq Jaffer, Jovana Knezevic et al.CVPR 2026 · 61 citations
- ProM3E: Probabilistic Masked MultiModal Embedding Model for EcologySrikumar Sastry, Subash Khanal, Aayush Dhakal, Jiayu Lin et al.CVPR 2026 · 1 citation
- MMLandmarks: a Cross-View Instance-Level Benchmark for Geo-Spatial UnderstandingOskar Kristoffersen, Alba Reinders Sánchez, Morten Rieger Hannemose, Anders Bjorholm Dahl et al.CVPR 2026
- Global and Local Entailment Learning for Natural World ImagerySrikumar Sastry, Aayush Dhakal, Eric Xing, Subash Khanal et al.ICCV 2025
Builds on20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- An Empirical Study of Training Self-Supervised Vision TransformersXinlei Chen, Saining Xie, Kaiming HeICCV 2021 · 2,340 citations
- SatMAE: Pre-training Transformers for Temporal and Multi-Spectral Satellite ImageryYezhen Cong, Samar Khanna, Chenlin Meng, Patrick Liu et al.NeurIPS 2022 · 707 citations
Related papers
- Combining Observational Data and Language for Species Range EstimationMax Hamilton, Christian Lange, Elijah Cole, Alexander Shepard et al.NeurIPS 2024 · 18 citations
- SatCLIP: Global, General-Purpose Location Embeddings with Satellite ImageryKonstantin Klemmer, Esther Rolf, Caleb Robinson, Lester Mackey et al.AAAI 2025 · 173 citations
- Benchmarking Representation Learning for Natural World Image CollectionsGrant Van Horn, Elijah Cole, Sara Beery, Kimberly Wilber et al.CVPR 2021
- CLIBD: Bridging Vision and Genomics for Biodiversity Monitoring at ScaleZeMing Gong, Austin T. Wang, Xiaoliang Huo, Joakim Bruslund Haurum et al.ICLR 2025
- Spectrally Distilled Representations Aligned with Instruction-Augmented LLMs for Satellite ImageryMinh Kha Do, Wei Xiang, Kang Han, Di Wu et al.CVPR 2026
