Remote Sensing Vision-Language Foundation Models without Annotations via Ground Remote Alignment
Utkarsh Mall, Cheng Perng Phoo, Meilin Kelsey Liu, Carl Vondrick, Bharath Hariharan, Kavita Bala
摘要
We introduce a method to train vision-language models for remote-sensing images without using any textual annotations. Our key insight is to use co-located internet imagery taken on the ground as an intermediary for connecting remote-sensing images and language. Specifically, we train an image encoder for remote sensing images to align with the image encoder of CLIP using a large amount of paired internet and satellite images. Our unsupervised approach enables the training of a first-of-its-kind large-scale vision language model (VLM) for remote sensing images at two different resolutions. We show that these VLMs enable zero-shot, open-vocabulary image classification, retrieval, segmentation and visual question answering for satellite images. On each of these tasks, our VLM trained without textual annotations outperforms existing VLMs trained with supervision, with gains of up to 20% for classification and 80% for segmentation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- Earth-Agent: Unlocking the Full Landscape of Earth Observation with AgentsPeilin Feng, Zhutao Lv, Junyan Ye, Xiaolei Wang 等ICLR 2026 · 被引用 49 次
- ZoomEarth: Active Perception for Ultra-High-Resolution Geospatial Vision-Language TasksRuixun Liu, Bowen Fu, Jiayi Song, Kaiyu Li 等CVPR 2026 · 被引用 19 次
- GOMAA-Geo: GOal Modality Agnostic Active Geo-localizationAnindya Sarkar, Srikumar Sastry, Aleksis Pirinen, Chongjie Zhang 等NeurIPS 2024 · 被引用 16 次
- Quality-Driven Curation of Remote Sensing Vision-Language Data via Learned Scoring ModelsDilxat Muhtar, Enzhuo Zhang, Zhenshi Li, Feng Gu 等NeurIPS 2025 · 被引用 16 次
- WildSAT: Learning Satellite Image Representations from Wildlife ObservationsRangel Daroya, Elijah Cole, Oisin Mac Aodha, Grant Van Horn 等ICCV 2025 · 被引用 4 次
它引用的顶会 Paper19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
相关 Paper
- SkyScript: A Large and Semantically Diverse Vision-Language Dataset for Remote SensingZhecheng Wang, Rajanie Prabha, Tianyuan Huang, Jiajun Wu 等AAAI 2024 · 被引用 167 次
- How Much Can CLIP Benefit Vision-and-Language Tasks?Sheng Shen, Liunian Harold Li, Hao Tan, Mohit Bansal 等ICLR 2022 · 被引用 503 次
- DiffCLIP: Few-shot Language-driven Multimodal ClassifierJiaqing Zhang, Mingxiang Cao, Xue Yang, Kai Jiang 等AAAI 2025 · 被引用 4 次
- Exploring Open-Vocabulary Semantic Segmentation from CLIP Vision Encoder Distillation OnlyJun Chen, Deyao Zhu, Guocheng Qian, Bernard Ghanem 等ICCV 2023 · 被引用 60 次
- Language-Driven Cross-Modal Classifier for Zero-Shot Multi-Label Image RecognitionYicheng Liu, Jie Wen, Chengliang Liu, Xiaozhao Fang 等ICML 2024 · 被引用 7 次
