Red-Teaming Text-to-Image Systems by Rule-based Preference Modeling
Yichuan Cao, Yibo Miao, Xiao-Shan Gao, Yinpeng Dong
摘要
Text-to-image (T2I) models raise ethical and safety concerns due to their potential to generate inappropriate or harmful images. Evaluating these models' security through red-teaming is vital, yet white-box approaches are limited by their need for internal access, complicating their use with closed-source models. Moreover, existing black-box methods often assume knowledge about the model's specific defense mechanisms, limiting their utility in real-world commercial API scenarios. A significant challenge is how to evade unknown and diverse defense mechanisms. To overcome this difficulty, we propose a novel Rule-based Preference modeling Guided Red-Teaming (RPG-RT), which iteratively employs LLM to modify prompts to query and leverages feedback from T2I systems for fine-tuning the LLM. RPG-RT treats the feedback from each iteration as a prior, enabling the LLM to dynamically adapt to unknown defense mechanisms. Given that the feedback is often labeled and coarse-grained, making it difficult to utilize directly, we further propose rule-based preference modeling, which employs a set of rules to evaluate desired or undesired feedback, facilitating finer-grained control over the LLM's dynamic adaptation process. Extensive experiments on nineteen T2I systems with varied safety mechanisms, three online commercial API services, and T2V models verify the superiority and practicality of our approach.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Jailbreaking on Text-to-Video Models via Scene Splitting StrategyWonjun Lee, Haon Park, Doehyeon Lee, Bumsub Ham 等ICLR 2026 · 被引用 10 次
- Time Is All It Takes: Spike-Retiming Attacks on Event-Driven Spiking Neural NetworksYi Yu, Qixin Zhang, Shuhan Ye, Xun Lin 等ICLR 2026 · 被引用 8 次
- Why Do Unlearnable Examples Work: A Novel Perspective of Mutual InformationYifan Zhu, Yibo Miao, Yinpeng Dong, Xiao-Shan GaoICLR 2026 · 被引用 3 次
- TEAR: Temporal-aware Automated Red-teaming for Text-to-Video ModelsJiaming He, Guanyu Hou, Hongwei Li, Zhicong Huang 等CVPR 2026 · 被引用 3 次
它引用的顶会 Paper34
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
相关 Paper
- Automated Red Teaming for Text-to-Image Models Through Feedback-Guided Prompt Iteration with Vision-Language ModelsWei Xu, Kangjie Chen, Jiawei Qiu, Yuyang Zhang 等ICCV 2025 · 被引用 3 次
- ART: Automatic Red-teaming for Text-to-Image Models to Protect Benign UsersGuanlin Li, Kangjie Chen, Shudong Zhang, Jie Zhang 等NeurIPS 2024 · 被引用 39 次
- FLIRT: Feedback Loop In-context Red TeamingNinareh Mehrabi, Palash Goyal, Christophe Dupuy, Qian Hu 等EMNLP 2024 · 被引用 8 次
- GenBreak: Red Teaming Text-to-Image Generation Using Large Language ModelsZilong Wang, Xiang Zheng, Xiaosen Wang, Bo Wang 等CVPR 2026
- DREAM: Scalable Red Teaming for Text-to-Image Generative Systems via Distribution ModelingBoheng Li, Junjie Wang, Yiming Li, Zhiyang Hu 等S&P 2026 · 被引用 9 次
