Modifier Unlocked: Jailbreaking Text-to-Image Models Through Prompts
Shuofeng Liu, Mengyao Ma, Minhui Xue, Guangdong Bai
Abstract
The unprecedented image generation capability of text-to-image models makes them double-edged swords. While these models allow users to create exquisite images through simple prompts, they also provide adversaries with opportunities to generate Not-Safe-for-Work (NSFW) content, referred to as the jailbreak attack. Despite built-in safety filters serving as a mitigation, their vulnerabilities and associated safety issues remain a significant concern. In this work, we propose MODX, the first modifier-based attack framework for jailbreaking text-to-image models. Modx leverages a heuristic algorithm with two heuristic functions (constraints) to identify modifiers that adjust the artistic genre to subtly introduce unsafe elements that drive the generated images towards NSFW. This approach takes advantage of the fact that filters are unlikely to reject images in certain styles or artistic forms, effectively inducing the models to generate NSFW content. We demonstrate the feasibility of modifier-based jailbreaking with a theoretical analysis, and provide experimental evidence of the effectiveness of MODX. Our results show that MODX outperforms existing methods in successfully achieving jailbreaking across four state-of-the-art text-to-image models. Moreover, we evaluate MODX across additional NSFW categories and on more models or model versions, demonstrating its strong scalability and generalization. Disclaimer: This paper contains NSFW language and imagery that could be offensive, distressing, and/or upsetting. Reader discretion is advised.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fa263fa6-220a-4bc1-94f2-0d0570a206b5Cited by top-tier papers5
- Odysseus: Jailbreaking Commercial Multimodal LLM-integrated Systems via Dual SteganographySongze Li, Jiameng Cheng, Yiming Li, Xiaojun Jia et al.NDSS 2026 · 9 citations
- FracFace: Breaking the Visual Clues - Fractal-Based Privacy-Preserving Face RecognitionWanying Dai, Beibei Li, Naipeng Dong, Guangdong Bai et al.NeurIPS 2025 · 7 citations
- SafeRoPE: Risk-specific Head-wise Embedding Rotation for Safe Generation in Rectified Flow TransformersXiang Yang, Feifei Li, Mi Zhang, Geng Hong et al.CVPR 2026 · 2 citations
- ReTrace: Reinforcement Learning-Guided Reconstruction Attacks on Machine UnlearningMengyao Ma, Shuofeng Liu, Minhui Xue, Surya Nepal et al.ICLR 2026
- Unified Safe In-context Image Generation in Multimodal Diffusion Transformers via Restricting Unsafe Information FlowsXiang Yang, Feifei Li, Mi Zhang, Geng Hong et al.ICML 2026
Builds on26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
Related papers
- SneakyPrompt: Jailbreaking Text-to-image Generative ModelsYuchen Yang, Bo Hui, Haolin Yuan, Neil Gong et al.S&P 2024 · 188 citations
- Universally Unfiltered and Unseen: Input-Agnostic Multimodal Jailbreaks against Text-to-Image Model SafeguardsSong Yan, Hui Wei, Jinlong Fei, Guoliang Yang et al.ACM MM 2025
- ColJailBreak: Collaborative Generation and Editing for Jailbreaking Text-to-Image Deep GenerationYizhuo Ma, Shanmin Pang, Qi Guo, Tianyu Wei et al.NeurIPS 2024 · 22 citations
- AdvPainting: Clean-text Jailbreaking Against Inpainting ModelsBingqian Zhou, Zhihao Wu, Yushi Cheng, Wenyuan XuACM MM 2025
- Multimodal Pragmatic Jailbreak on Text-to-image ModelsTong Liu, Zhixin Lai, Jiawen Wang, Gengyuan Zhang et al.ACL 2025
