S-CLIP: Semi-supervised Vision-Language Learning using Few Specialist Captions
Sangwoo Mo, Minkyu Kim, Kyungmin Lee, Jinwoo Shin
摘要
Vision-language models, such as contrastive language-image pre-training (CLIP), have demonstrated impressive results in natural image domains. However, these models often struggle when applied to specialized domains like remote sensing, and adapting to such domains is challenging due to the limited number of image-text pairs available for training. To address this, we propose S-CLIP, a semi-supervised learning method for training CLIP that utilizes additional unpaired images. S-CLIP employs two pseudo-labeling strategies specifically designed for contrastive learning and the language modality. The caption-level pseudo-label is given by a combination of captions of paired images, obtained by solving an optimal transport problem between unpaired and paired images. The keyword-level pseudo-label is given by a keyword in the caption of the nearest paired image, trained through partial label learning that assumes a candidate set of labels for supervision instead of the exact one. By combining these objectives, S-CLIP significantly enhances the training of CLIP using only a few image-text pairs, as demonstrated in various specialist domains, including remote sensing, fashion, scientific figures, and comics. For instance, S-CLIP improves CLIP by 10% for zero-shot classification and 4% for image-text retrieval on the remote sensing benchmark, matching the performance of supervised CLIP while using three times fewer image-text pairs. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Knowledge Composition using Task Vectors with Learned Anisotropic ScalingFrederic Z. Zhang, Paul Albert, Cristian Rodriguez Opazo, Anton van den Hengel 等NeurIPS 2024 · 被引用 43 次
- S-MolSearch: 3D Semi-supervised Contrastive Learning for Bioactive Molecule SearchGengmo Zhou, Zhen Wang, Feng Yu, Guolin Ke 等NeurIPS 2024 · 被引用 8 次
- DiffCLIP: Few-shot Language-driven Multimodal ClassifierJiaqing Zhang, Mingxiang Cao, Xue Yang, Kai Jiang 等AAAI 2025 · 被引用 4 次
- Learning Shared Representations from Unpaired DataAmitai Yacobi, Nir Ben-Ari, Ronen Talmon, Uri ShahamNeurIPS 2025 · 被引用 3 次
- AffordMatcher: Affordance Learning in 3D Scenes from Visual SignifiersNghia Vu, Tuong Do, Khang Nguyen, Baoru Huang 等CVPR 2026 · 被引用 2 次
它引用的顶会 Paper36
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna 等NeurIPS 2020 · 被引用 7,049 次
相关 Paper
- Semi-Supervised CLIP Adaptation by Enforcing Semantic and Trapezoidal ConsistencyKai Gan, Bo Ye, Min-Ling Zhang, Tong WeiICLR 2025
- SuperCLIP: CLIP with Simple Classification SupervisionWeiheng Zhao, Zilong Huang, Jiashi Feng, Xinggang WangNeurIPS 2025 · 被引用 6 次
- Toward Modality Gap: Vision Prototype Learning for Weakly-supervised Semantic Segmentation with CLIPZhongxing Xu, Feilong Tang, Zhe Chen, Yingxue Su 等AAAI 2025 · 被引用 23 次
- RankCLIP: Ranking-Consistent Language-Image PretrainingYiming Zhang, Zhuokai Zhao, Zhaorun Chen, Zhili Feng 等ICCV 2025 · 被引用 1 次
- Domain-Controlled Prompt LearningQinglong Cao, Zhengqin Xu, Yuntian Chen, Chao Ma 等AAAI 2024 · 被引用 38 次
