CSP: Self-Supervised Contrastive Spatial Pre-Training for Geospatial-Visual Representations
Gengchen Mai, Ni Lao, Yutong He, Jiaming Song, Stefano Ermon
Abstract
Geo-tagged images are publicly available in large quantities, whereas labels such as object classes are rather scarce and expensive to collect. Meanwhile, contrastive learning has achieved tremendous success in various natural image and language tasks with limited labeled data. However, existing methods fail to fully leverage geospatial information, which can be paramount to distinguishing objects that are visually similar. To directly leverage the abundant geospatial information associated with images in pre-training, fine-tuning, and inference stages, we present Contrastive Spatial Pre-Training (CSP), a selfsupervised learning framework for geo-tagged images. We use a dual-encoder to separately encode the images and their corresponding geo-locations, and use contrastive objectives to learn effective location representations from images, which can be transferred to downstream supervised tasks such as image classification. Experiments show that CSP can improve model performance on both iNat2018 and fMoW dataset. Especially, on iNat2018, CSP significantly boosts the model performance with 10-34% relative improvement with various labeled training data sampling ratios 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c54d87cd-e7cb-4841-bd7f-43c5672380e0Cited by top-tier papers16
- CROMA: Remote Sensing Representations with Contrastive Radar-Optical Masked AutoencodersAnthony Fuller, Koreen Millard, James R. GreenNeurIPS 2023 · 245 citations
- SatCLIP: Global, General-Purpose Location Embeddings with Satellite ImageryKonstantin Klemmer, Esther Rolf, Caleb Robinson, Lester Mackey et al.AAAI 2025 · 173 citations
- GeoLLM: Extracting Geospatial Knowledge from Large Language ModelsRohin Manvi, Samar Khanna, Gengchen Mai, Marshall Burke et al.ICLR 2024 · 104 citations
- Geographic Location Encoding with Spherical Harmonics and Sinusoidal Representation NetworksMarc Rußwurm, Konstantin Klemmer, Esther Rolf, Robin Zbinden et al.ICLR 2024 · 66 citations
- Measuring the Intrinsic Dimension of Earth RepresentationsArjun Rao, Marc Rußwurm, Konstantin Klemmer, Esther RolfICLR 2026 · 13 citations
Builds on14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
Related papers
- Geography-Aware Self-Supervised LearningKumar Ayush, Burak Uzkent, Chenlin Meng, Kumar Tanmay et al.ICCV 2021 · 304 citations
- Spatially Consistent Representation LearningByungseok Roh, Wuhyun Shin, Ildoo Kim, Sungwoong KimCVPR 2021
- From Coarse to Fine: A Matching and Alignment Framework for Unsupervised Cross-View Geo-LocalizationXueyi Wang, Lele Zhang, Zheng Fan, Yang Liu et al.AAAI 2025 · 12 citations
- Unsupervised Object-Level Representation Learning from Scene ImagesJiahao Xie, Xiaohang Zhan, Ziwei Liu, Yew Soon Ong et al.NeurIPS 2021 · 93 citations
- Unleashing the Power of Contrastive Self-Supervised Visual Models via Contrast-Regularized Fine-TuningYifan Zhang, Bryan Hooi, Dapeng Hu, Jian Liang et al.NeurIPS 2021 · 82 citations
