Stabilizing Modality Gap & Lowering Gradient Norms Improve Zero-Shot Adversarial Robustness of VLMs
Junhao Dong, Piotr Koniusz, Xinghua Qu, Yew-Soon Ong
Abstract
Contemporary Vision-Language Models (VLMs) such as CLIP offer an attractive zero-shot classification functionality facilitated by large-scale vision-language pre-training. However, they remain vulnerable to adversarial attacks, a critical security threat in realistic deployment. Adversarially robust fine-tuning provides generalizable robustness on new datasets while preserving natural performance by fine-tuning the pre-trained models. Fine-tuning robust CLIP typically relies on adversaries generated solely from the vision branch. However, this singular focus on the vision modality, coupled with static text prompts used as fixed category prototypes, limits the robustness achieved through dual-modality fine-tuning. We observe for CLIP fine-tuning that zero-shot adversarial robustness improves when we (i) stabilize the modality gap (a phenomenon where image and text features occupy different feature space regions) and (ii) lower/stabilize gradient norms. Both these steps enjoy further improvement of robustness if one fine-tunes with both visual and text adversaries. For both modalities, we leverage (i) the maximization of an effective rank of features and (ii) noise modulation of features. We show that maximizing the effective rank helps lower and stabilize the modality gap over adversaries with varying perturbation radii. The noise modulation of features, achieved by the so-called count sketching, lowers/stabilizes gradient norms. We outperform the state of the art on 15 datasets. We provide the first insights into the effects of modality gap & gradient norms in VLM fine-tuning.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 42a2ece1-c84a-42c4-a8ee-e4ec9be8e064Cited by top-tier papers22
- Fit the Distribution: Cross-Image/Prompt Adversarial Attacks on Multimodal Large Language ModelsHai Yan, Haijian Ma, Xiaowen Cai, Daizong Liu et al.NeurIPS 2025 · 21 citations
- Enhancing CLIP Robustness via Cross-Modality AlignmentXingyu Zhu, Beier Zhu, Shuo Wang, Kesen Zhao et al.NeurIPS 2025 · 17 citations
- CogniVerse: Revolutionizing Multi-Modal Retrieval-Augmented Generation with Cognitive Reflection and Geometric ReasoningXiang Fang, Wanlong Fang, Changshuo WangCVPR 2026 · 17 citations
- Machine Unlearning via Task Simplex ArithmeticJunhao Dong, Hao Zhu, Yifei Zhang, Xinghua Qu et al.NeurIPS 2025 · 10 citations
- Towards Building Model/Prompt-Transferable Attackers against Large Vision-Language ModelsXiaowen Cai, Daizong Liu, Xiaoye Qu, Xiang Fang et al.NeurIPS 2025 · 8 citations
Related papers
- Understanding Zero-shot Adversarial Robustness for Large-Scale ModelsChengzhi Mao, Scott Geng, Junfeng Yang, Xin Wang et al.ICLR 2023 · 10 citations
- Quality Text, Robust Vision: The Role of Language in Enhancing Visual Robustness of Vision-Language ModelsFuta Waseda, Saku Sugawara, Isao EchizenACM MM 2025 · 2 citations
- TAPT: Test-Time Adversarial Prompt Tuning for Robust Inference in Vision-Language ModelsXin Wang, Kai Chen, Jiaming Zhang, Jingjing Chen et al.CVPR 2025
- On the Zero-shot Adversarial Robustness of Vision-Language Models: A Truly Zero-shot and Training-free ApproachBaoshun Tong, Hanjiang Lai, Yan Pan, Jian YinCVPR 2025
- Pre-Trained Model Guided Fine-Tuning for Zero-Shot Adversarial RobustnessSibo Wang, Jie Zhang, Zheng Yuan, Shiguang ShanCVPR 2024
