Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models
Christian Schlarmann, Naman Deep Singh, Francesco Croce, Matthias Hein
Abstract
Multi-modal foundation models like OpenFlamingo, LLaVA, and GPT-4 are increasingly used for various real-world tasks. Prior work has shown that these models are highly vulnerable to adversarial attacks on the vision modality. These attacks can be leveraged to spread fake information or defraud users, and thus pose a significant risk, which makes the robustness of large multi-modal foundation models a pressing problem. The CLIP model, or one of its variants, is used as a frozen vision encoder in many large vision-language models (LVLMs), e.g. LLaVA and OpenFlamingo. We propose an unsupervised adversarial fine-tuning scheme to obtain a robust CLIP vision encoder, which yields robustness on all vision down-stream tasks (LVLMs, zero-shot classification) that rely on CLIP. In particular, we show that stealth-attacks on users of LVLMs by a malicious third party providing manipulated images are no longer possible once one replaces the original CLIP model with our robust one. No retraining or fine-tuning of the down-stream LVLMs is required. The code and robust models are available at https://github.com/chs20/RobustVLM
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ade9437c-368e-4732-8ddf-948dfd02a97dCited by top-tier papers75
- Text-Guided Attention is All You Need for Zero-Shot Robustness in Vision-Language ModelsLu Yu, Haiyang Zhang, Changsheng XuNeurIPS 2024 · 29 citations
- Enhancing CLIP Robustness via Cross-Modality AlignmentXingyu Zhu, Beier Zhu, Shuo Wang, Kesen Zhao et al.NeurIPS 2025 · 17 citations
- MIP against Agent: Malicious Image Patches Hijacking Multimodal OS AgentsLukas Aichberger, Alasdair Paren, Guohao Li, Philip H. S. Torr et al.NeurIPS 2025 · 12 citations
- How Dark Patterns Manipulate Web AgentsPhil Cuvin, Hao Zhu, Diyi YangICLR 2026 · 9 citations
- AdPO: Enhancing the Adversarial Robustness of Large Vision-Language Models with Preference OptimizationChaohu Liu, Tianyi Gui, Yu Liu, Linli XuICLR 2026 · 9 citations
Builds on29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
Related papers
- On Evaluating Adversarial Robustness of Large Vision-Language ModelsYunqing Zhao, Tianyu Pang, Chao Du, Xiao Yang et al.NeurIPS 2023 · 404 citations
- SLADE: Shielding against Dual Exploits in Large Vision-Language ModelsMd. Zarif Hossain, Ahmed ImteajCVPR 2025
- TAPT: Test-Time Adversarial Prompt Tuning for Robust Inference in Vision-Language ModelsXin Wang, Kai Chen, Jiaming Zhang, Jingjing Chen et al.CVPR 2025
- Image Hijacks: Adversarial Images can Control Generative Models at RuntimeLuke Bailey, Euan Ong, Stuart Russell, Scott EmmonsICML 2024 · 171 citations
- Robustness in Both Domains: CLIP Needs a Robust Text EncoderElías Abad-Rocamora, Christian Schlarmann, Naman Deep Singh, Yongtao Wu et al.NeurIPS 2025 · 4 citations
