Ideator: Jailbreaking and Benchmarking Large Vision-Language Models Using Themselves
Ruofan Wang, Juncheng Li, Yixu Wang, Bo Wang, Xiaosen Wang, Yan Teng, Yingchun Wang, Xingjun Ma, Yu-Gang Jiang
摘要
As large Vision-Language Models (VLMs) gain prominence, ensuring their safe deployment has become critical. Recent studies have explored VLM robustness against jailbreak at-tacks-techniques that exploit model vulnerabilities to elicit harmful outputs. However, the limited availability of diverse multimodal data has constrained current approaches to rely heavily on adversarial or manually crafted images derived from harmful text datasets, which often lack effectiveness and diversity across different contexts. In this paper, we propose IDEATOR, a novel jailbreak method that autonomously generates malicious image-text pairs for black-box jailbreak attacks. IDEATOR is grounded in the insight that VLMs themselves could serve as powerful red team models for generating multimodal jailbreak prompts. Specifically, IDEATOR leverages a VLM to create targeted jailbreak texts and pairs them with jailbreak images generated by a state-of-the-art diffusion model. Extensive experiments demonstrate IDEATOR's high effectiveness and transferability, achieving a 94% attack success rate (ASR) in jailbreaking MiniGPT-4 with an average of only 5.34 queries, and high ASRs of , and 75 % when transferred to LLaVA, InstructBLIP, and Chameleon, respectively. Building on IDEATOR's strong transferability and automated process, we introduce the VLJailbreakBench, a safety benchmark comprising 3,654 multimodal jailbreak samples. Our benchmark results on 11 recently released VLMs reveal significant gaps in safety alignment. For instance, our challenge set achieves ASRs of 46.31% on GPT40 and 19.65% on Claude-3.5-Sonnet, underscoring the urgent need for stronger defenses. Disclaimer: This paper contains content that may be disturbing or offensive.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Attention! Your Vision Language Model Could Be Maliciously ManipulatedXiaosen Wang, Shaokang Wang, Zhijin Ge, Yuyang Luo 等NeurIPS 2025 · 被引用 13 次
- Evolving Contextual Safety in Multi-Modal Large Language Models via Inference-Time Self-Reflective MemoryCe Zhang, Jinxi He, Junyi He, Katia Sycara 等CVPR 2026 · 被引用 5 次
- Toward Universal and Transferable Jailbreak Attacks on Vision-Language ModelsKaiyuan Cui, Yige Li, Yutao Wu, Xingjun Ma 等ICLR 2026 · 被引用 4 次
- Dictionary-Aligned Concept Control for Safeguarding Multimodal LLMsJinqi Luo, Jinyu Yang, Tal Neiman, Lei Fan 等CVPR 2026 · 被引用 1 次
- Are Vision-Language Models Safe in the Wild? A Meme-Based Benchmark StudyDongGeon Lee, Joonwon Jang, Jihae Jeong, Hwanjo YuEMNLP 2025 · 被引用 1 次
它引用的顶会 Paper15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
相关 Paper
- Jailbreaking Multimodal Large Language Models via Shuffle InconsistencyShiji Zhao, Ranjie Duan, Fengxiang Wang, Chi Chen 等ICCV 2025 · 被引用 8 次
- Exploring Visual Vulnerabilities via Multi-Loss Adversarial Search for Jailbreaking Vision-Language ModelsShuyang Hao, Bryan Hooi, Jun Liu, Kai-Wei Chang 等CVPR 2025
- LLMs Caught in the Crossfire: Malware Requests and Jailbreak ChallengesHaoyang Li, Huan Gao, Zhiyuan Zhao, Zhiyu Lin 等ACL 2025
- MultiBreak: A Scalable and Diverse Multi-turn Jailbreak Benchmark for Evaluating LLM SafetyJialin Song, Xiaodong Liu, Weiwei Yang, Wuyang Chen 等ICML 2026 · 被引用 5 次
- White-box Multimodal Jailbreaks Against Large Vision-Language ModelsRuofan Wang, Xingjun Ma, Hanxu Zhou, Chuanjun Ji 等ACM MM 2024 · 被引用 22 次
