Enhancing CLIP Robustness via Cross-Modality Alignment
Xingyu Zhu, Beier Zhu, Shuo Wang, Kesen Zhao, Hanwang Zhang
摘要
Vision-language models (VLMs) such as CLIP demonstrate strong generalization in zero-shot classification but remain highly vulnerable to adversarial perturbations. Existing methods primarily focus on adversarial fine-tuning or prompt optimization, they often overlook the gaps in CLIP's encoded features, which is shown as the text and image features lie far apart from each other. This misalignment is significantly amplified under adversarial perturbations, leading to severe degradation in classification performance. To address this problem, we propose CrOss-modaLity Alignment, dubbed COLA, an optimal transport-based framework that explicitly addresses adversarial misalignment by restoring both global image-text alignment and local structural consistency in the feature space. (1) COLA first projects adversarial image embeddings onto a subspace spanned by class text features, effectively filtering out non-semantic distortions while preserving discriminative information. (2) It then models images and texts as discrete distributions over multiple augmented views and refines their alignment via OT, with the subspace projection seamlessly integrated into the cost computation. This design ensures stable cross-modal alignment even under adversarial conditions. COLA is trainingfree and compatible with existing fine-tuned models. Extensive evaluations across 14 zero-shot classification benchmarks demonstrate the effectiveness of COLA, especially with an average improvement of 6.7% on ImageNet and its variants under PGD adversarial attacks, while maintaining high accuracy on clean samples.
Recent efforts to improve the adversarial robustness of VLMs can be broadly categorized into three directions: adversarial training [4, 62], which fine-tunes models with perturbed samples; prompt tuning [31,63], which optimizes text input templates to resist attacks, and test-time defenses [21,2,59], which modify inputs or predictions on the fly. While these methods offer promising improvements, they suffer from high computational overhead or introduce substantial inference latency. More critically, they overlook a central issue: the misalignment between image and text modalities [70,16]. This misalignment stems from CLIP's global matching paradigm, where the model is trained to align entire image embeddings with sentence-level textual embeddings. As shown in Figure 1(a), the text † Corresponding author 39th Conference on Neural Information Processing Systems (NeurIPS 2025).
2 Related Work Adversarial robustness in VLMs. Adversarial robustness remains a fundamental challenge, as small, imperceptible perturbations can mislead model predictions [51,7]. A common defense is adversarial training (AT) [33,62,47], which improves robustness but incurs high computational cost [50,56]. Recent test-time defenses-such as generative purification [37,61] and optimization-based methods [57, 35]-offer alternatives but often fail under adaptive attacks [12]. Hedge Defense (HD) [57], for example, perturbs inputs by maximizing cross-entropy loss, but requires an adversarially trained model. Meanwhile, several works explore CLIP's [45] robustness, noting its natural tendency to deflect attacks via counteractive perturbations in latent space. To further improve performance, researchers have applied adversarial fine-tuning [36,54] and prompt tuning with frozen weights [31,63].
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- On the Generalization of SFT: A Reinforcement Learning Perspective with Reward RectificationYongliang Wu, Yizhou Zhou, Ziheng Zhou, Yingzhe Peng 等ICLR 2026 · 被引用 130 次
- Look Carefully: Adaptive Visual Reinforcements in Multimodal Large Language Models for Hallucination MitigationXingyu Zhu, Kesen Zhao, Liang Yi, Shuo Wang 等ICLR 2026 · 被引用 9 次
- Streaming Drag-Oriented Interactive Video Manipulation: Drag Anything, Anytime!Junbao Zhou, Yuan Zhou, Kesen Zhao, Qingshan Xu 等ICLR 2026 · 被引用 7 次
- Revisiting Robustness for LLM Safety Alignment via Selective Geometry ControlYonghui Yang, Wenjian Tao, Jilong Liu, Xingyu Zhu 等ICML 2026 · 被引用 4 次
- GuardAlign: Test-time Safety Alignment in Multimodal Large Language ModelsXingyu Zhu, Beier Zhu, Junfeng Fang, Shuo Wang 等ICLR 2026 · 被引用 2 次
它引用的顶会 Paper37
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 被引用 9,786 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution GeneralizationDan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath 等ICCV 2021 · 被引用 2,294 次
- Fast is better than free: Revisiting adversarial trainingEric Wong, Leslie Rice, J. Zico KolterICLR 2020 · 被引用 1,352 次
相关 Paper
- TAPT: Test-Time Adversarial Prompt Tuning for Robust Inference in Vision-Language ModelsXin Wang, Kai Chen, Jiaming Zhang, Jingjing Chen 等CVPR 2025
- Robustifying Vision-Language Models via Test-Time Prompt AdaptationXingyu Zhu, Huanshen Wu, Shuo Wang, Beier Zhu 等ICML 2026 · 被引用 1 次
- Identifying Robust Neural Pathways: Few-Shot Adversarial Mask Tuning for Vision-Language ModelsWonjeong Choi, Sejong Ryu, Jungmoon Lee, Dong-Jun Han 等ICLR 2026
- Understanding Zero-shot Adversarial Robustness for Large-Scale ModelsChengzhi Mao, Scott Geng, Junfeng Yang, Xin Wang 等ICLR 2023 · 被引用 10 次
- Self-Calibrated Consistency can Fight Back for Adversarial Robustness in Vision-Language ModelsJiaxiang Liu, Jiawei Du, Xiao Liu, Shangyang Li 等ICML 2026 · 被引用 2 次
