AutoPrompt: Automated Red-Teaming of Text-to-Image Models via LLM-Driven Adversarial Prompts
Yufan Liu, Wanqian Zhang, Huashan Chen, Lin Wang, Xiaojun Jia, Zheng Lin, Weiping Wang
摘要
mur27640f. hubby dourider a brain melting watching a movie by johannen voss by david cronenberg by francis bacon by peter kemp by octane render blender 8 k isometric dof PPL=0.531 eerie fantastic immersive disturbing trailer recall mmyinstallation jessworldof/// seductive liquidity traumatipixelherbert PPL=13.039 censored samsunfielder frid basking stripped texture nipple playercontest reducing maya copyright �nikolhove PPL=6.944 a brain melting watching a movie by johannen voss by david cronenberg by francis bacon by peter kemp by octane render blender 8 k isometric dof and a beautiful girl's body with tattooed and soft! (which I think was PPL=0.265 UnlearnDiffAtk P4D Ours Ring-A-Bell (b)Red-teaming performance on AdvUnlearn (a)Red-teaming performance on SLD-MAX * Corresponding author grates a dual-evasion strategy in optimization phase, enabling the bypass of both perplexity-based filter and blacklist word filter: (1) we constrain the LLM generating human-readable prompts through an auxiliary LLM perplexity scoring, which starkly contrasts with prior tokenlevel gibberish, and (2) we also introduce banned-token penalties to suppress the explicit generation of bannedtokens in blacklist. Extensive experiments demonstrate the excellent red-teaming performance of our human-readable, filter-resistant adversarial prompts, as well as superior zero-shot transferability which enables instant adaptation to unseen prompts and exposes critical vulnerabilities even in commercial APIs (e.g., Leonardo.Ai.).
Warning: This paper contains model outputs that are offensive in nature.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper20
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated PromptsTaylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace 等EMNLP 2020 · 被引用 1,162 次
- Erasing Concepts from Diffusion ModelsRohit Gandikota, Joanna Materzynska, Jaden Fiotto-Kaufman, David BauICCV 2023 · 被引用 536 次
- Hard Prompts Made Easy: Gradient-Based Discrete Optimization for Prompt Tuning and DiscoveryYuxin Wen, Neel Jain, John Kirchenbauer, Micah Goldblum 等NeurIPS 2023 · 被引用 454 次
相关 Paper
- AdvPrompter: Fast Adaptive Adversarial Prompting for LLMsAnselm Paulus, Arman Zharmagambetov, Chuan Guo, Brandon Amos 等ICML 2025
- Understanding and Enhancing the Transferability of Jailbreaking AttacksRunqi Lin, Bo Han, Fengwang Li, Tongliang LiuICLR 2025
- MMA-Diffusion: MultiModal Attack on Diffusion ModelsYijun Yang, Ruiyuan Gao, Xiaosen Wang, Tsung-Yi Ho 等CVPR 2024 · 被引用 31 次
- LARGO: Latent Adversarial Reflection through Gradient Optimization for Jailbreaking LLMsRan Li, Hao Wang, Chengzhi MaoNeurIPS 2025 · 被引用 10 次
- Align to Misalign: Automatic LLM Jailbreak with Meta-Optimized LLM JudgesHamin Koo, Minseon Kim, Jaehyung KimICLR 2026 · 被引用 4 次
