SoftCLIP: Softer Cross-Modal Alignment Makes CLIP Stronger
Yuting Gao, Jinfeng Liu, Zihan Xu, Tong Wu, Enwei Zhang, Ke Li, Jie Yang, Wei Liu, Xing Sun
摘要
During the preceding biennium, vision-language pre-training has achieved noteworthy success on several downstream tasks. Nevertheless, acquiring high-quality image-text pairs, where the pairs are entirely exclusive of each other, remains a challenging task, and noise exists in the commonly used datasets. To address this issue, we propose SoftCLIP, a novel approach that relaxes the strict one-to-one constraint and achieves a soft cross-modal alignment by introducing a softened target, which is generated from the fine-grained intra-modal self-similarity. The intra-modal guidance is indicative to enable two pairs have some local similarities and model many-to-many relationships between the two modalities. Besides, since the positive still dominates in the softened target distribution, we disentangle the negatives in the distribution to further boost the relation alignment with the negatives in the cross-modal learning. Extensive experiments demonstrate the effectiveness of SoftCLIP. In particular, on ImageNet zero-shot classification task, using CC3M/CC12M as pre-training dataset, SoftCLIP brings a top-1 accuracy improvement of 6.8%/7.2% over the CLIP baseline.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper27
- Cross-Modal and Uni-Modal Soft-Label Alignment for Image-Text RetrievalHailang Huang, Zhijie Nie, Ziqiao Wang, Ziyu ShangAAAI 2024 · 被引用 47 次
- Scalable Incomplete Multi-View Clustering with Structure AlignmentYi Wen, Siwei Wang, Ke Liang, Weixuan Liang 等ACM MM 2023 · 被引用 36 次
- CWCL: Cross-Modal Transfer with Continuously Weighted Contrastive LossRakshith Sharma Srinivasa, Jaejin Cho, Chouchang Yang, Yashas Malur Saidutta 等NeurIPS 2023 · 被引用 25 次
- SynBrain: Enhancing Visual-to-fMRI Synthesis via Probabilistic Representation LearningWeijian Mai, Jiamin Wu, Yu Zhu, Zhouheng Yao 等NeurIPS 2025 · 被引用 11 次
- ModalChorus: Visual Probing and Alignment of Multi-Modal Embeddings via Modal Fusion MapYilin Ye, Shishi Xiao, Xingchen Zeng, Wei ZengIEEE VIS 2024 · 被引用 8 次
它引用的顶会 Paper8
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty 等NeurIPS 2021 · 被引用 2,985 次
- FILIP: Fine-grained Interactive Language-Image Pre-TrainingLewei Yao, Runhui Huang, Lu Hou, Guansong Lu 等ICLR 2022 · 被引用 827 次
- Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training ParadigmYangguang Li, Feng Liang, Lichen Zhao, Yufeng Cui 等ICLR 2022 · 被引用 565 次
相关 Paper
- RankCLIP: Ranking-Consistent Language-Image PretrainingYiming Zhang, Zhuokai Zhao, Zhaorun Chen, Zhili Feng 等ICCV 2025 · 被引用 1 次
- PyramidCLIP: Hierarchical Feature Alignment for Vision-language Model PretrainingYuting Gao, Jinfeng Liu, Zihan Xu, Jun Zhang 等NeurIPS 2022 · 被引用 168 次
- Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality InversionMarco Mistretta, Alberto Baldrati, Lorenzo Agnolucci, Marco Bertini 等ICLR 2025
- Semi-Supervised CLIP Adaptation by Enforcing Semantic and Trapezoidal ConsistencyKai Gan, Bo Ye, Min-Ling Zhang, Tong WeiICLR 2025
- Mitigate the Gap: Improving Cross-Modal Alignment in CLIPSedigheh Eslami, Gerard de MeloICLR 2025 · 被引用 1 次
