PLA: Prompt Learning Attack Against Text-To-Image Generative Models
Xinqi Lyu, Yihao Liu, Yanjie Li, Bin Xiao
摘要
Text-to-Image (T2I) models have gained widespread adoption across various applications. Despite the success, the potential misuse of T2I models poses significant risks of generating Not-Safe-For-Work (NSFW) content. To investigate the vulnerability of T2I models, this paper delves into adversarial attacks to bypass the safety mechanisms under black-box settings. Most previous methods rely on word substitution to search adversarial prompts. Due to limited search space, this leads to suboptimal performance compared to gradient-based training. However, black-box settings present unique challenges to training gradient-driven attack methods, since there is no access to the internal architecture and parameters of T2I models. To facilitate the learning of adversarial prompts in black-box settings, we propose a novel prompt learning attack framework (), where insightful gradient-based training tailored to blackbox T2I models is designed by utilizing multimodal similarities. Experiments show that our new method can effectively attack the safety mechanisms of black-box T2I models including prompt filters and post-hoc safety checkers with a high success rate compared to state-of-the-art methods. Warning: This paper may contain offensive modelgenerated content.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- LOMIA: Label-Only Membership Inference Attacks against Pre-trained Large Vision-Language ModelsYihao Liu, Xinqi Lyu, Dong Wang, Yanjie Li 等NeurIPS 2025 · 被引用 3 次
- Cert-LAS: Toward Certified Model Ownership Verification for Text-to-Image Diffusion Models via Layer-Adaptive SmoothingLeyi Qi, Yiming Li, Siyuan Liang, Zhengzhong Tu 等ICML 2026 · 被引用 1 次
- FeatureFool: Zero-Query Fooling of Video Models via Feature MapDuoxun Tang, Xi Xiao, Guangwu Hu, Kangkang Sun 等CVPR 2026 · 被引用 1 次
- Breaking Multimodal LLM Safety via Video-Driven PromptingDong Wang, XIANGYU HE, Xinqi Lyu, Bin XiaoCVPR 2026
- Red-teaming Retrieval-Augmented Diffusion Models via Poisoning Knowledge BasesXinqi Lyu, Yihao Liu, Dong Wang, Bin XiaoCVPR 2026
它引用的顶会 Paper18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann 等ICLR 2024 · 被引用 4,569 次
相关 Paper
- MacPrompt: Maraconic-Guided Jailbreak Against Text-to-Image ModelsXi Ye, Yiwen Liu, Lina Wang, Run Wang 等AAAI 2026
- Transstratal Adversarial Attack: Compromising Multi-Layered Defenses in Text-to-Image ModelsChunlong Xie, Kangjie Chen, Shangwei Guo, Shudong Zhang 等NeurIPS 2025 · 被引用 1 次
- AdvI2I: Adversarial Image Attack on Image-to-Image Diffusion ModelsYaopei Zeng, Yuanpu Cao, Bochuan Cao, Yurui Chang 等ICML 2025
- Modifier Unlocked: Jailbreaking Text-to-Image Models Through PromptsShuofeng Liu, Mengyao Ma, Minhui Xue, Guangdong BaiS&P 2025
- SneakyPrompt: Jailbreaking Text-to-image Generative ModelsYuchen Yang, Bo Hui, Haolin Yuan, Neil Gong 等S&P 2024 · 被引用 188 次
