Perception-Guided Jailbreak Against Text-to-Image Models
Yihao Huang, Le Liang, Tianlin Li, Xiaojun Jia, Run Wang, Weikai Miao, Geguang Pu, Yang Liu
Abstract
In recent years, Text-to-Image (T2I) models have garnered significant attention due to their remarkable advancements. However, security concerns have emerged due to their potential to generate inappropriate or Not-Safe-For-Work (NSFW) images. In this paper, inspired by the observation that texts with different semantics can lead to similar human perceptions, we propose an LLM-driven perception-guided jailbreak method, termed PGJ. It is a black-box jailbreak method that requires no specific T2I model (model-free) and generates highly natural attack prompts. Specifically, we propose identifying a safe phrase that is similar in human perception yet inconsistent in text semantics with the target unsafe word and using it as a substitution. The experiments conducted on six open-source models and commercial online services with thousands of prompts have verified the effectiveness of PGJ. Warning: This paper contains NSFW and disturbing imagery, including adult, violent, and illegal-related contentious content. We have masked images deemed unsafe. However, reader discretion is advised.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a716196e-1d85-46f0-961e-9704be7be8caCited by top-tier papers26
- Adversarial Attacks against Closed-Source MLLMs via Feature Optimal AlignmentXiaojun Jia, Sensen Gao, Simeng Qin, Tianyu Pang et al.NeurIPS 2025 · 49 citations
- Heuristic-Induced Multimodal Risk Distribution Jailbreak Attack for Multimodal Large Language ModelsTeng Ma, Xiaojun Jia, Ranjie Duan, Xinfeng Li et al.ICCV 2025 · 35 citations
- T2V-OptJail: Discrete Prompt Optimization for Text-to-Video Jailbreak AttacksJiayang Liu, Siyuan Liang, Shiqian Zhao, Rong-Cheng Tu et al.NeurIPS 2025 · 18 citations
- JailbreakDiffBench: A Comprehensive Benchmark for Jailbreaking Diffusion ModelsXiaolong Jin, Zixuan Weng, Hanxi Guo, Chenlong Yin et al.ICCV 2025 · 13 citations
- DREAM: Scalable Red Teaming for Text-to-Image Generative Systems via Distribution ModelingBoheng Li, Junjie Wang, Yiming Li, Zhiyang Hu et al.S&P 2026 · 9 citations
Builds on24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
Related papers
- Modifier Unlocked: Jailbreaking Text-to-Image Models Through PromptsShuofeng Liu, Mengyao Ma, Minhui Xue, Guangdong BaiS&P 2025
- Reason2Attack: Jailbreaking Text-to-Image Models via LLM ReasoningChenyu Zhang, Lanjun Wang, Yiwen Ma, Wenhui Li et al.AAAI 2026 · 7 citations
- ColJailBreak: Collaborative Generation and Editing for Jailbreaking Text-to-Image Deep GenerationYizhuo Ma, Shanmin Pang, Qi Guo, Tianyu Wei et al.NeurIPS 2024 · 22 citations
- Fuzz-Testing Meets LLM-Based Agents: An Automated and Efficient Framework for Jailbreaking Text-to-Image Generation ModelsYingkai Dong, Xiangtao Meng, Ning Yu, Zheng Li et al.S&P 2025
- PLA: Prompt Learning Attack Against Text-To-Image Generative ModelsXinqi Lyu, Yihao Liu, Yanjie Li, Bin XiaoICCV 2025 · 10 citations
