Pre-training CLIP against Data Poisoning with Optimal Transport-based Matching and Alignment
Tong Zhang, Kuofeng Gao, Jiawang Bai, Leo Yu Zhang, Xin Yin, Zonghui Wang, Shouling Ji, Wenzhi Chen
摘要
Recent studies have shown that Contrastive Language-Image Pre-training (CLIP) models are threatened by targeted data poisoning and backdoor attacks due to massive training image-caption pairs crawled from the Internet. Previous defense methods correct poisoned image-caption pairs by matching a new caption for each image. However, the matching process relies solely on the global representations of images and captions, overlooking fine-grained features of visual and textual features. It may introduce incorrect image-caption pairs and harm the CLIP pre-training. To address their limitations, we propose an Optimal Transport-based framework to reconstruct image-caption pairs, named OTCCLIP. We propose a new optimal transport-based distance measure between fine-grained visual and textual feature sets and re-assign new captions based on the proposed optimal transport distance. Additionally, to further reduce the negative impact of mismatched pairs, we encourage the inter- and intra-modality fine-grained alignment by employing optimal transport-based objective functions. Our experiments demonstrate that OTCCLIP can successfully decrease the attack success rates of poisoning attacks. Also, compared to previous methods, OTCCLIP significantly improves CLIP's zero-shot and linear probing performance trained on poisoned datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Grounding Language with Vision: A Conditional Mutual Information Calibrated Decoding Strategy for Reducing Hallucinations in LVLMsHao Fang, Changle Zhou, Jiawei Kong, Kuofeng Gao 等NeurIPS 2025 · 被引用 25 次
- HALoRA: Low-Rank Adaptation with Hierarchical Budget Allocation for Efficient Vision-Language AlignmentLetian Zhang, Guanghao Meng, Xudong Ren, Jinpeng WangAAAI 2026
它引用的顶会 Paper23
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- Conditional Prompt Learning for Vision-Language ModelsKaiyang Zhou, Jingkang Yang, Chen Change Loy, Ziwei LiuCVPR 2022 · 被引用 1,438 次
- FILIP: Fine-grained Interactive Language-Image Pre-TrainingLewei Yao, Runhui Huang, Lu Hou, Guansong Lu 等ICLR 2022 · 被引用 827 次
- Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training ParadigmYangguang Li, Feng Liang, Lichen Zhao, Yufeng Cui 等ICLR 2022 · 被引用 565 次
- Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual ConceptsYan Zeng, Xinsong Zhang, Hang LiICML 2022 · 被引用 371 次
相关 Paper
- Better Safe than Sorry: Pre-training CLIP against Targeted Data Poisoning and Backdoor AttacksWenhan Yang, Jingdong Gao, Baharan MirzasoleimanICML 2024 · 被引用 21 次
- Robust Contrastive Language-Image Pretraining against Data Poisoning and Backdoor AttacksWenhan Yang, Jingdong Gao, Baharan MirzasoleimanNeurIPS 2023 · 被引用 51 次
- ToxicTextCLIP: Text-Based Poisoning and Backdoor Attacks on CLIP Pre-trainingXin Yao, Haiyang Zhao, Yimin Chen, Jiawei Guo 等NeurIPS 2025 · 被引用 5 次
- OT-CLIP: Understanding and Generalizing CLIP via Optimal TransportLiangliang Shi, Jack Fan, Junchi YanICML 2024 · 被引用 11 次
- Detecting Backdoor Samples in Contrastive Language Image PretrainingHanxun Huang, Sarah Monazam Erfani, Yige Li, Xingjun Ma 等ICLR 2025
