Automated Red Teaming for Text-to-Image Models Through Feedback-Guided Prompt Iteration with Vision-Language Models
Wei Xu, Kangjie Chen, Jiawei Qiu, Yuyang Zhang, Run Wang, Jin Mao, Tianwei Zhang, Lina Wang
摘要
Text-to-image models have achieved remarkable progress in generating high-quality images from textual prompts, yet their potential for misuse like generating unsafe content remains a critical concern. Existing safety mechanisms, such as filtering and fine-tuning, remain insufficient in preventing vulnerabilities exposed by adversarial prompts. To systematically evaluate these weaknesses, we propose an automated red-teaming framework, Feedback-Guided Prompt Iteration (FGPI), which utilizes a Vision-Language Model (VLM) as the red-teaming agent following a feedback-guide-rewrite paradigm for iterative prompt optimization. The red-teaming VLM analyzes prompt-image pairs based on evaluation results, provides feedback and modification strategies to enhance adversarial effectiveness while preserving safety constraints, and iteratively improves prompts. To enable this functionality, we construct a multi-turn conversational VQA dataset with over 6,000 instances, covering seven attack types and facilitating the fine-tuning of the red-teaming VLM. Extensive experiments demonstrate the effectiveness of our approach, achieving over 90% attack success rate within five iterations while maintaining prompt stealthiness and safety. The experiments also validate the adaptability, diversity, transferability, and explainability of FGPI. The source code and dataset are available at https://github.com/Weiww-Xu/FGPI.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- TEAR: Temporal-aware Automated Red-teaming for Text-to-Video ModelsJiaming He, Guanyu Hou, Hongwei Li, Zhicong Huang 等CVPR 2026 · 被引用 3 次
- MacPrompt: Maraconic-Guided Jailbreak Against Text-to-Image ModelsXi Ye, Yiwen Liu, Lina Wang, Run Wang 等AAAI 2026
它引用的顶会 Paper18
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
相关 Paper
- Iterative Prompt Refinement for Safer Text-to-Image GenerationJinwoo Jeon, JunHyeok Oh, Hayeong Lee, Byung-Jun LeeEMNLP 2025
- TRUST-VLM: Thorough Red-Teaming for Uncovering Safety Threats in Vision-Language ModelsKangjie Chen, Muyang Li, Guanlin Li, Shudong Zhang 等ICML 2025
- ART: Automatic Red-teaming for Text-to-Image Models to Protect Benign UsersGuanlin Li, Kangjie Chen, Shudong Zhang, Jie Zhang 等NeurIPS 2024 · 被引用 39 次
- DREAM: Scalable Red Teaming for Text-to-Image Generative Systems via Distribution ModelingBoheng Li, Junjie Wang, Yiming Li, Zhiyang Hu 等S&P 2026 · 被引用 9 次
- FLIRT: Feedback Loop In-context Red TeamingNinareh Mehrabi, Palash Goyal, Christophe Dupuy, Qian Hu 等EMNLP 2024 · 被引用 8 次
