WildSAT: Learning Satellite Image Representations from Wildlife Observations
Rangel Daroya, Elijah Cole, Oisin Mac Aodha, Grant Van Horn, Subhransu Maji
摘要
Species distributions encode valuable ecological and environmental information, yet their potential for guiding representation learning in remote sensing remains underexplored. We introduce WildSAT, which pairs satellite images with millions of geo-tagged wildlife observations readily-available on citizen science platforms. WildSAT employs a contrastive learning approach that jointly leverages satellite images, species occurrence maps, and textual habitat descriptions to train or fine-tune models. This approach significantly improves performance on diverse satellite image recognition tasks, outperforming both ImageNet-pretrained models and satellite-specific baselines. Additionally, by aligning visual and textual information, WildSAT enables zero-shot retrieval, allowing users to search geographic locations based on textual descriptions. WildSAT surpasses recent cross-modal learning methods, including approaches that align satellite images with ground imagery or wildlife photos, demonstrating the advantages of our approach. Finally, we analyze the impact of key design choices and highlight the broad applicability of WildSAT to remote sensing and biodiversity monitoring.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- TESSERA: Temporal Embeddings of Surface Spectra for Earth Representation and AnalysisZhengpeng Feng, Clement Atzberger, Sadiq Jaffer, Jovana Knezevic 等CVPR 2026 · 被引用 61 次
- ProM3E: Probabilistic Masked MultiModal Embedding Model for EcologySrikumar Sastry, Subash Khanal, Aayush Dhakal, Jiayu Lin 等CVPR 2026 · 被引用 1 次
- MMLandmarks: a Cross-View Instance-Level Benchmark for Geo-Spatial UnderstandingOskar Kristoffersen, Alba Reinders Sánchez, Morten Rieger Hannemose, Anders Bjorholm Dahl 等CVPR 2026
- Global and Local Entailment Learning for Natural World ImagerySrikumar Sastry, Aayush Dhakal, Eric Xing, Subash Khanal 等ICCV 2025
它引用的顶会 Paper20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- An Empirical Study of Training Self-Supervised Vision TransformersXinlei Chen, Saining Xie, Kaiming HeICCV 2021 · 被引用 2,340 次
- SatMAE: Pre-training Transformers for Temporal and Multi-Spectral Satellite ImageryYezhen Cong, Samar Khanna, Chenlin Meng, Patrick Liu 等NeurIPS 2022 · 被引用 707 次
相关 Paper
- Combining Observational Data and Language for Species Range EstimationMax Hamilton, Christian Lange, Elijah Cole, Alexander Shepard 等NeurIPS 2024 · 被引用 18 次
- SatCLIP: Global, General-Purpose Location Embeddings with Satellite ImageryKonstantin Klemmer, Esther Rolf, Caleb Robinson, Lester Mackey 等AAAI 2025 · 被引用 173 次
- Benchmarking Representation Learning for Natural World Image CollectionsGrant Van Horn, Elijah Cole, Sara Beery, Kimberly Wilber 等CVPR 2021
- CLIBD: Bridging Vision and Genomics for Biodiversity Monitoring at ScaleZeMing Gong, Austin T. Wang, Xiaoliang Huo, Joakim Bruslund Haurum 等ICLR 2025
- Spectrally Distilled Representations Aligned with Instruction-Augmented LLMs for Satellite ImageryMinh Kha Do, Wei Xiang, Kang Han, Di Wu 等CVPR 2026
