Fuzz-Testing Meets LLM-Based Agents: An Automated and Efficient Framework for Jailbreaking Text-to-Image Generation Models
Yingkai Dong, Xiangtao Meng, Ning Yu, Zheng Li, Shanqing Guo
摘要
Text-to-image (T2I) generative models have revolutionized content creation by transforming textual descriptions into high-quality images. However, these models are vulnerable to jailbreaking attacks, where carefully crafted prompts bypass safety mechanisms to produce unsafe content. While researchers have developed various jailbreak attacks to expose this risk, these methods face significant limitations, including impractical access requirements, easily detectable unnatural prompts, restricted search spaces, and high query demands on the target system. In this paper, we propose JailFuzzer, a novel fuzzing framework driven by large language model (LLM) agents, designed to efficiently generate natural and semantically meaningful jailbreak prompts in a black-box setting. Specifically, JailFuzzer employs fuzz-testing principles with three components: a seed pool for initial and jailbreak prompts, a guided mutation engine for generating meaningful variations, and an oracle function to evaluate jailbreak success. Furthermore, we construct the guided mutation engine and oracle function by LLM-based agents, which further ensures efficiency and adaptability in black-box settings. Extensive experiments demonstrate that JailFuzzer has significant advantages in jailbreaking T2I models. It generates natural and semantically coherent prompts, reducing the likelihood of detection by traditional defenses. Additionally, it achieves a high success rate in jailbreak attacks with minimal query overhead, outperforming existing methods across all key metrics. This study underscores the need for stronger safety mechanisms in generative models and provides a foundation for future research on defending against sophisticated jailbreaking attacks. JailFuzzer is open-source and available at this repository: https://github.com/YingkaiD/JailFuzzer.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- T2V-OptJail: Discrete Prompt Optimization for Text-to-Video Jailbreak AttacksJiayang Liu, Siyuan Liang, Shiqian Zhao, Rong-Cheng Tu 等NeurIPS 2025 · 被引用 18 次
- JailbreakDiffBench: A Comprehensive Benchmark for Jailbreaking Diffusion ModelsXiaolong Jin, Zixuan Weng, Hanxi Guo, Chenlong Yin 等ICCV 2025 · 被引用 13 次
- DREAM: Scalable Red Teaming for Text-to-Image Generative Systems via Distribution ModelingBoheng Li, Junjie Wang, Yiming Li, Zhiyang Hu 等S&P 2026 · 被引用 9 次
- Red-Teaming Text-to-Image Systems by Rule-based Preference ModelingYichuan Cao, Yibo Miao, Xiao-Shan Gao, Yinpeng DongNeurIPS 2025 · 被引用 8 次
- RunawayEvil: Jailbreaking the Image-to-Video Generative Modelsyueming lyu, Rufan Qian, Yueming Lyu, Qinglong Liu 等CVPR 2026 · 被引用 7 次
它引用的顶会 Paper19
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann 等ICLR 2024 · 被引用 4,569 次
- Improving Factuality and Reasoning in Language Models through Multiagent DebateYilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum 等ICML 2024 · 被引用 1,562 次
相关 Paper
- OrchJail: Jailbreaking Tool-Calling Text-to-Image Agents by Orchestration-Guided FuzzingJianming Chen, Yawen Wang, Junjie Wang, Zhe Liu 等ICML 2026
- LLM-Fuzzer: Scaling Assessment of Large Language Model JailbreaksJiahao Yu, Xingwei Lin, Zheng Yu, Xinyu XingUSENIX Security 2024 · 被引用 83 次
- PAPILLON: Efficient and Stealthy Fuzz Testing-Powered Jailbreaks for LLMsXueluan Gong, Mingzhe Li, Yilin Zhang, Fengyuan Ran 等USENIX Security 2025
- Efficient LLM-Jailbreaking via Multimodal-LLM JailbreakHaoxuan Ji, Zheng Lin, Zhenxing Niu, Xinbo Gao 等AAAI 2026 · 被引用 4 次
- Stand on The Shoulders of Giants: Building JailExpert from Previous Attack ExperienceXi Wang, Songlei Jian, Shasha Li, Xiaopeng Li 等EMNLP 2025 · 被引用 1 次
