AutoPrompt: Automated Red-Teaming of Text-to-Image Models via LLM-Driven Adversarial Prompts
Yufan Liu, Wanqian Zhang, Huashan Chen, Lin Wang, Xiaojun Jia, Zheng Lin, Weiping Wang
Abstract
mur27640f. hubby dourider a brain melting watching a movie by johannen voss by david cronenberg by francis bacon by peter kemp by octane render blender 8 k isometric dof PPL=0.531 eerie fantastic immersive disturbing trailer recall mmyinstallation jessworldof/// seductive liquidity traumatipixelherbert PPL=13.039 censored samsunfielder frid basking stripped texture nipple playercontest reducing maya copyright �nikolhove PPL=6.944 a brain melting watching a movie by johannen voss by david cronenberg by francis bacon by peter kemp by octane render blender 8 k isometric dof and a beautiful girl's body with tattooed and soft! (which I think was PPL=0.265 UnlearnDiffAtk P4D Ours Ring-A-Bell (b)Red-teaming performance on AdvUnlearn (a)Red-teaming performance on SLD-MAX * Corresponding author grates a dual-evasion strategy in optimization phase, enabling the bypass of both perplexity-based filter and blacklist word filter: (1) we constrain the LLM generating human-readable prompts through an auxiliary LLM perplexity scoring, which starkly contrasts with prior tokenlevel gibberish, and (2) we also introduce banned-token penalties to suppress the explicit generation of bannedtokens in blacklist. Extensive experiments demonstrate the excellent red-teaming performance of our human-readable, filter-resistant adversarial prompts, as well as superior zero-shot transferability which enables instant adaptation to unseen prompts and exposes critical vulnerabilities even in commercial APIs (e.g., Leonardo.Ai.).
Warning: This paper contains model outputs that are offensive in nature.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on20
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated PromptsTaylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace et al.EMNLP 2020 · 1,162 citations
- Erasing Concepts from Diffusion ModelsRohit Gandikota, Joanna Materzynska, Jaden Fiotto-Kaufman, David BauICCV 2023 · 536 citations
- Hard Prompts Made Easy: Gradient-Based Discrete Optimization for Prompt Tuning and DiscoveryYuxin Wen, Neel Jain, John Kirchenbauer, Micah Goldblum et al.NeurIPS 2023 · 454 citations
Related papers
- AdvPrompter: Fast Adaptive Adversarial Prompting for LLMsAnselm Paulus, Arman Zharmagambetov, Chuan Guo, Brandon Amos et al.ICML 2025
- Understanding and Enhancing the Transferability of Jailbreaking AttacksRunqi Lin, Bo Han, Fengwang Li, Tongliang LiuICLR 2025
- MMA-Diffusion: MultiModal Attack on Diffusion ModelsYijun Yang, Ruiyuan Gao, Xiaosen Wang, Tsung-Yi Ho et al.CVPR 2024 · 31 citations
- LARGO: Latent Adversarial Reflection through Gradient Optimization for Jailbreaking LLMsRan Li, Hao Wang, Chengzhi MaoNeurIPS 2025 · 10 citations
- Align to Misalign: Automatic LLM Jailbreak with Meta-Optimized LLM JudgesHamin Koo, Minseon Kim, Jaehyung KimICLR 2026 · 4 citations
