FLIRT: Feedback Loop In-context Red Teaming
Ninareh Mehrabi, Palash Goyal, Christophe Dupuy, Qian Hu, Shalini Ghosh, Richard S. Zemel, Kai-Wei Chang, Aram Galstyan, Rahul Gupta
Abstract
Warning: this paper contains content that may be inappropriate or offensive. As generative models become available for public use in various applications, testing and analyzing vulnerabilities of these models has become a priority. In this work, we propose an automatic red teaming framework that evaluates a given black-box model and exposes its vulnerabilities against unsafe and inappropriate content generation. Our framework uses incontext learning in a feedback loop to red team models and trigger them into unsafe content generation. In particular, taking text-to-image models as target models, we explore different feedback mechanisms to automatically learn effective and diverse adversarial prompts. Our experiments demonstrate that even with enhanced safety features, Stable Diffusion (SD) models are vulnerable to our adversarial prompts, raising concerns on their robustness in practical uses. Furthermore, we demonstrate that the proposed framework is effective for red teaming text-to-text models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ceb4d760-8f9f-4dc7-8605-4614b11f274aCited by top-tier papers14
- ART: Automatic Red-teaming for Text-to-Image Models to Protect Benign UsersGuanlin Li, Kangjie Chen, Shudong Zhang, Jie Zhang et al.NeurIPS 2024 · 39 citations
- Unveiling the Basin-Like Loss Landscape in Large Language ModelsHuanran Chen, Zeming Wei, Yao Huang, Yichi Zhang et al.ICLR 2026 · 14 citations
- AdversaFlow: Visual Red Teaming for Large Language Models with Multi-Level Adversarial FlowDazhen Deng, Chuhan Zhang, Huawei Zheng, Yuwen Pu et al.IEEE VIS 2024 · 14 citations
- MoGU: A Framework for Enhancing Safety of LLMs While Preserving Their UsabilityYanrui Du, Sendong Zhao, Danyang Zhao, Ming Ma et al.NeurIPS 2024 · 13 citations
- DREAM: Scalable Red Teaming for Text-to-Image Generative Systems via Distribution ModelingBoheng Li, Junjie Wang, Yiming Li, Zhiyang Hu et al.S&P 2026 · 9 citations
Builds on6
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Self-Instruct: Aligning Language Models with Self-Generated InstructionsYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu et al.ACL 2023 · 540 citations
- Red Teaming Language Models with Language ModelsEthan Perez, Saffron Huang, H. Francis Song, Trevor Cai et al.EMNLP 2022 · 239 citations
- Neural Path Hunter: Reducing Hallucination in Dialogue Systems via Path GroundingNouha Dziri, Andrea Madotto, Osmar Zaïane, Avishek Joey BoseEMNLP 2021 · 74 citations
- Query-Efficient Black-Box Red Teaming via Bayesian OptimizationDeokjae Lee, JunYeong Lee, Jung-Woo Ha, Jin-Hwa Kim et al.ACL 2023 · 5 citations
Related papers
- GenBreak: Red Teaming Text-to-Image Generation Using Large Language ModelsZilong Wang, Xiang Zheng, Xiaosen Wang, Bo Wang et al.CVPR 2026
- Prompting4Debugging: Red-Teaming Text-to-Image Diffusion Models by Finding Problematic PromptsZhi-Yi Chin, Chieh-Ming Jiang, Ching-Chun Huang, Pin-Yu Chen et al.ICML 2024 · 155 citations
- Ring-A-Bell! How Reliable are Concept Removal Methods For Diffusion Models?Yu-Lin Tsai, Chia-Yi Hsu, Chulin Xie, Chih-Hsun Lin et al.ICLR 2024 · 207 citations
- PLA: Prompt Learning Attack Against Text-To-Image Generative ModelsXinqi Lyu, Yihao Liu, Yanjie Li, Bin XiaoICCV 2025 · 10 citations
- SneakyPrompt: Jailbreaking Text-to-image Generative ModelsYuchen Yang, Bo Hui, Haolin Yuan, Neil Gong et al.S&P 2024 · 188 citations
