Failures to Find Transferable Image Jailbreaks Between Vision-Language Models
Rylan Schaeffer, Dan Valentine, Luke Bailey, James Chua, Cristóbal Eyzaguirre, Zane Durante, Joe Benton, Brando Miranda, Henry Sleight, Tony Tong Wang, John Hughes, Rajashree Agrawal
摘要
The integration of new modalities into frontier AI systems increases the possibility such systems can be adversarially manipulated in undesirable ways. In this work, we focus on a popular class of vision-language models (VLMs) that generate text conditioned on visual and textual inputs. We conducted a large-scale empirical study to assess the transferability of gradient-based universal image "jailbreaks" using a diverse set of over 40 open-parameter VLMs, including 18 new VLMs that we publicly release. We find that transferable gradient-based image jailbreaks are extremely difficult to obtain. When an image jailbreak is optimized against a single VLM or against an ensemble of VLMs, the image successfully jailbreaks the attacked VLM(s), but exhibits little-to-no transfer to any other VLMs; transfer is not affected by whether the attacked and target VLMs possess matching vision backbones or language models, whether the language model underwent instructionfollowing and/or safety-alignment training, or other factors. Only two settings display partial transfer: between identically-pretrained and identically-initialized VLMs with slightly different VLM training data, and between different training checkpoints of a single VLM. Leveraging these results, we demonstrate that transfer can be significantly improved against a specific target VLM by attacking larger ensembles of "highly-similar" VLMs. These results stand in stark contrast to existing evidence of universal and transferable text jailbreaks against language models and transferable adversarial attacks against image classifiers, suggesting that VLMs may be more robust to gradient-based transfer attacks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Speech-Audio Compositional Attacks on Multimodal LLMs and Their Defense with SALMONN-GuardYudong Yang, Xuezhen Zhang, Zhifeng Han, Siyin Wang 等ICML 2026 · 被引用 13 次
- FORCE: Transferable Visual Jailbreaking Attacks via Feature Over-Reliance CorrEctionRunqi Lin, Alasdair Paren, Suqin Yuan, Muyang Li 等CVPR 2026 · 被引用 13 次
- MIP against Agent: Malicious Image Patches Hijacking Multimodal OS AgentsLukas Aichberger, Alasdair Paren, Guohao Li, Philip H. S. Torr 等NeurIPS 2025 · 被引用 12 次
- Towards Building Model/Prompt-Transferable Attackers against Large Vision-Language ModelsXiaowen Cai, Daizong Liu, Xiaoye Qu, Xiang Fang 等NeurIPS 2025 · 被引用 8 次
- Jailbreak Transferability Emerges from Shared RepresentationsRico Angell, Jannik Brinkmann, He HeICLR 2026 · 被引用 5 次
它引用的顶会 Paper32
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
相关 Paper
- Toward Universal and Transferable Jailbreak Attacks on Vision-Language ModelsKaiyuan Cui, Yige Li, Yutao Wu, Xingjun Ma 等ICLR 2026 · 被引用 4 次
- White-box Multimodal Jailbreaks Against Large Vision-Language ModelsRuofan Wang, Xingjun Ma, Hanxu Zhou, Chuanjun Ji 等ACM MM 2024 · 被引用 22 次
- Jailbreak in pieces: Compositional Adversarial Attacks on Multi-Modal Language ModelsErfan Shayegani, Yue Dong, Nael B. Abu-GhazalehICLR 2024 · 被引用 271 次
- FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual PromptsYichen Gong, Delong Ran, Jinyuan Liu, Conglei Wang 等AAAI 2025 · 被引用 350 次
- Jailbreaking Vision-Language Models Through the Visual ModalityAharon Azulay, Jan Dubiński, Zhuoyun Li, Atharv Mittal 等ICML 2026 · 被引用 3 次
