Enhancing CLIP Robustness via Cross-Modality Alignment
Xingyu Zhu, Beier Zhu, Shuo Wang, Kesen Zhao, Hanwang Zhang
Abstract
Vision-language models (VLMs) such as CLIP demonstrate strong generalization in zero-shot classification but remain highly vulnerable to adversarial perturbations. Existing methods primarily focus on adversarial fine-tuning or prompt optimization, they often overlook the gaps in CLIP's encoded features, which is shown as the text and image features lie far apart from each other. This misalignment is significantly amplified under adversarial perturbations, leading to severe degradation in classification performance. To address this problem, we propose CrOss-modaLity Alignment, dubbed COLA, an optimal transport-based framework that explicitly addresses adversarial misalignment by restoring both global image-text alignment and local structural consistency in the feature space. (1) COLA first projects adversarial image embeddings onto a subspace spanned by class text features, effectively filtering out non-semantic distortions while preserving discriminative information. (2) It then models images and texts as discrete distributions over multiple augmented views and refines their alignment via OT, with the subspace projection seamlessly integrated into the cost computation. This design ensures stable cross-modal alignment even under adversarial conditions. COLA is trainingfree and compatible with existing fine-tuned models. Extensive evaluations across 14 zero-shot classification benchmarks demonstrate the effectiveness of COLA, especially with an average improvement of 6.7% on ImageNet and its variants under PGD adversarial attacks, while maintaining high accuracy on clean samples.
Recent efforts to improve the adversarial robustness of VLMs can be broadly categorized into three directions: adversarial training [4, 62], which fine-tunes models with perturbed samples; prompt tuning [31,63], which optimizes text input templates to resist attacks, and test-time defenses [21,2,59], which modify inputs or predictions on the fly. While these methods offer promising improvements, they suffer from high computational overhead or introduce substantial inference latency. More critically, they overlook a central issue: the misalignment between image and text modalities [70,16]. This misalignment stems from CLIP's global matching paradigm, where the model is trained to align entire image embeddings with sentence-level textual embeddings. As shown in Figure 1(a), the text † Corresponding author 39th Conference on Neural Information Processing Systems (NeurIPS 2025).
2 Related Work Adversarial robustness in VLMs. Adversarial robustness remains a fundamental challenge, as small, imperceptible perturbations can mislead model predictions [51,7]. A common defense is adversarial training (AT) [33,62,47], which improves robustness but incurs high computational cost [50,56]. Recent test-time defenses-such as generative purification [37,61] and optimization-based methods [57, 35]-offer alternatives but often fail under adaptive attacks [12]. Hedge Defense (HD) [57], for example, perturbs inputs by maximizing cross-entropy loss, but requires an adversarially trained model. Meanwhile, several works explore CLIP's [45] robustness, noting its natural tendency to deflect attacks via counteractive perturbations in latent space. To further improve performance, researchers have applied adversarial fine-tuning [36,54] and prompt tuning with frozen weights [31,63].
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 423a5aeb-544c-45bd-8b19-dffe62c064cbCited by top-tier papers13
- On the Generalization of SFT: A Reinforcement Learning Perspective with Reward RectificationYongliang Wu, Yizhou Zhou, Ziheng Zhou, Yingzhe Peng et al.ICLR 2026 · 130 citations
- Look Carefully: Adaptive Visual Reinforcements in Multimodal Large Language Models for Hallucination MitigationXingyu Zhu, Kesen Zhao, Liang Yi, Shuo Wang et al.ICLR 2026 · 9 citations
- Streaming Drag-Oriented Interactive Video Manipulation: Drag Anything, Anytime!Junbao Zhou, Yuan Zhou, Kesen Zhao, Qingshan Xu et al.ICLR 2026 · 7 citations
- Revisiting Robustness for LLM Safety Alignment via Selective Geometry ControlYonghui Yang, Wenjian Tao, Jilong Liu, Xingyu Zhu et al.ICML 2026 · 4 citations
- GuardAlign: Test-time Safety Alignment in Multimodal Large Language ModelsXingyu Zhu, Beier Zhu, Junfeng Fang, Shuo Wang et al.ICLR 2026 · 2 citations
Builds on37
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution GeneralizationDan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath et al.ICCV 2021 · 2,294 citations
- Fast is better than free: Revisiting adversarial trainingEric Wong, Leslie Rice, J. Zico KolterICLR 2020 · 1,352 citations
Related papers
- TAPT: Test-Time Adversarial Prompt Tuning for Robust Inference in Vision-Language ModelsXin Wang, Kai Chen, Jiaming Zhang, Jingjing Chen et al.CVPR 2025
- Robustifying Vision-Language Models via Test-Time Prompt AdaptationXingyu Zhu, Huanshen Wu, Shuo Wang, Beier Zhu et al.ICML 2026 · 1 citation
- Identifying Robust Neural Pathways: Few-Shot Adversarial Mask Tuning for Vision-Language ModelsWonjeong Choi, Sejong Ryu, Jungmoon Lee, Dong-Jun Han et al.ICLR 2026
- Understanding Zero-shot Adversarial Robustness for Large-Scale ModelsChengzhi Mao, Scott Geng, Junfeng Yang, Xin Wang et al.ICLR 2023 · 10 citations
- Self-Calibrated Consistency can Fight Back for Adversarial Robustness in Vision-Language ModelsJiaxiang Liu, Jiawei Du, Xiao Liu, Shangyang Li et al.ICML 2026 · 2 citations
