Moderator: Moderating Text-to-Image Diffusion Models through Fine-grained Context-based Policies
Peiran Wang, Qiyu Li, Longxuan Yu, Ziyao Wang, Ang Li, Haojian Jin
摘要
We present Moderator, a policy-based model management system that allows administrators to specify fine-grained content moderation policies and modify the weights of a text-to-image (TTI) model to make it significantly more challenging for users to produce images that violate the policies. In contrast to existing general-purpose model editing techniques, which unlearn concepts without considering the associated contexts, Moderator allows admins to specify what content should be moderated, under which context, how it should be moderated, and why moderation is necessary. Given a set of policies, Moderator first prompts the original model to generate images that need to be moderated, then uses these self-generated images to reverse fine-tune the model to compute task vectors for moderation and finally negates the original model with the task vectors to decrease its performance in generating moderated content. We evaluated Moderator with 14 participants to play the role of admins and found they could quickly learn and author policies to pass unit tests in approximately 2.29 policy iterations. Our experiment with 32 stable diffusion users suggested that Moderator can prevent 65% of users from generating moderated content under 15 attempts and require the remaining users an average of 8.3 times more attempts to generate undesired content. CCS CONCEPTS • Security and privacy → Human and societal aspects of security and privacy; • Social and professional topics → Computing / technology policy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- VidLaDA: Bidirectional Diffusion Large Language Models for Efficient Video UnderstandingZhihao He, Tieyuan Chen, Kangyu Wang, Ziran Qin 等ICML 2026 · 被引用 3 次
- Selective Fine-Tuning for Targeted and Robust Concept UnlearningMansi Mansi, Avinash Kori, Francesca Toni, Soteris DemetriouCCS 2026 · 被引用 2 次
- Attacks on Approximate Caches in Text-to-Image Diffusion ModelsDesen Sun, Shuncheng Jie, Sihang LiuUSENIX Security 2026 · 被引用 1 次
- SafeGuider: Robust and Practical Content Safety Control for Text-to-Image ModelsPeigui Qi, Kunsheng Tang, Wenbo Zhou, Weiming Zhang 等CCS 2025
- CHAIRO: Contextual Hierarchical Analogical Induction and Reasoning Optimization for LLMsHaotian Lu, Yuchen Mou, Bingzhe WuACL 2026
它引用的顶会 Paper35
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann 等ICLR 2024 · 被引用 4,569 次
- Certified Data Removal from Machine Learning ModelsChuan Guo, Tom Goldstein, Awni Y. Hannun, Laurens van der MaatenICML 2020 · 被引用 633 次
相关 Paper
- Meta-Unlearning on Diffusion Models: Preventing Relearning Unlearned ConceptsHongcheng Gao, Tianyu Pang, Chao Du, Taihang Hu 等ICCV 2025 · 被引用 4 次
- Growth Inhibitors for Suppressing Inappropriate Image Concepts in Diffusion ModelsDie Chen, Zhiwen Li, Mingyuan Fan, Cen Chen 等ICLR 2025
- "I Cannot Write This Because It Violates Our Content Policy": Understanding Content Moderation Policies and User Experiences in Generative AI ProductsLan Gao, Oscar Chen, Rachel Lee, Nick Feamster 等USENIX Security 2025
- GuardT2I: Defending Text-to-Image Models from Adversarial PromptsYijun Yang, Ruiyuan Gao, Xiao Yang, Jianyuan Zhong 等NeurIPS 2024 · 被引用 74 次
- Ablating Concepts in Text-to-Image Diffusion ModelsNupur Kumari, Bingliang Zhang, Sheng-Yu Wang, Eli Shechtman 等ICCV 2023 · 被引用 327 次
