AEGIS: Adversarial Target-Guided Retention-Data-Free Robust Concept Erasure from Diffusion Models
Fengpeng Li, Kemou Li, Qizhou Wang, Bo Han, Jiantao Zhou
摘要
Concept erasure helps stop diffusion models (DMs) from generating harmful content; but current methods face robustness-retention trade-off. Robustness means the model fine-tuned by concept erasure methods resists reactivation of erased concepts, even under semantically related prompts. Retention means unrelated concepts are preserved so the model’s overall utility stays intact. Both are critical for concept erasure in practice, yet addressing them simultaneously is challenging, as existing works typically improve one factor while sacrificing the other. Prior work typically strengthens one while degrading the other—e.g., mapping a single erased prompt to a fixed safe target leaves class-level remnants exploitable by prompt attacks, whereas retention-oriented schemes underperform against adaptive adversaries. This paper introduces Adversarial Erasure with Gradient-Informed Synergy (AEGIS), a retention-data-free framework that advances both robustness and retention. First, AEGIS replaces handpicked targets with an Adversarial Erasure Target (AET) optimized to approximate the semantic center of the erased concept class. By aligning the model’s prediction on the erased prompt to an AET-derived target in the shared text–image space, AEGIS increases predicted-noise distances not just for the instance but for semantically related variants, substantially hardening the DMs against state-of-the-art adversarial prompt attacks. Second, AEGIS preserves utility without auxiliary data via Gradient Regularization Projection (GRP), a conflict-aware gradient rectification that selectively projects away the destructive component of the retention update only when it opposes the erasure direction. This directional, data-free projection mitigates interference between erasure and retention, avoiding dataset bias and accidental relearning. Extensive experiments show that AEGIS markedly reduces attack success rates across various concepts while maintaining or improving FID/CLIP versus advanced baselines, effectively pushing beyond the prevailing robustness–retention trade-off. The source code is in https://github.com/Feng-peng-Li/AEGIS.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- LLM Unlearning with LLM BeliefsKemou Li, Qizhou Wang, Yue Wang, Fengpeng Li 等ICLR 2026 · 被引用 20 次
- When Robots Obey the Patch: Universal Transferable Patch Attacks on Vision-Language-Action ModelsHui Lu, Yi Yu, Yiming Yang, Chenyu Yi 等CVPR 2026 · 被引用 12 次
- Editprint: General Digital Image Forensics via Editing Fingerprint with Self-Augmentation TrainingHaiwei Wu, Kemou Li, Yuanman Li, Jiantao ZhouCVPR 2026
- Zero-shot Detection of AI-Generated Image via RAW-RGB AlignmentHaiwei Wu, Fengpeng Li, Zhilin Tu, Yuanman Li 等CVPR 2026
它引用的顶会 Paper34
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine 等NeurIPS 2020 · 被引用 2,261 次
- Erasing Concepts from Diffusion ModelsRohit Gandikota, Joanna Materzynska, Jaden Fiotto-Kaufman, David BauICCV 2023 · 被引用 536 次
- Hard Prompts Made Easy: Gradient-Based Discrete Optimization for Prompt Tuning and DiscoveryYuxin Wen, Neel Jain, John Kirchenbauer, Micah Goldblum 等NeurIPS 2023 · 被引用 454 次
相关 Paper
- Fantastic Targets for Concept Erasure in Diffusion Models and Where To Find ThemAnh Tuan Bui, Thuy-Trang Vu, Long Tung Vuong, Trung Le 等ICLR 2025
- Erasing Undesirable Concepts in Diffusion Models with Adversarial PreservationAnh Bui, Tung-Long Vuong, Khanh Doan, Trung Le 等NeurIPS 2024 · 被引用 55 次
- One-dimensional Adapter to Rule Them All: Concepts, Diffusion Models and Erasing ApplicationsMengyao Lyu, Yuhong Yang, Haiwen Hong, Hui Chen 等CVPR 2024 · 被引用 16 次
- GrOCE : Graph-Guided Online Concept Erasure for Text-to-Image Diffusion ModelsNing Han, Zhenyu Ge, Feng Han, Yuhua Sun 等CVPR 2026 · 被引用 3 次
- Defensive Unlearning with Adversarial Training for Robust Concept Erasure in Diffusion ModelsYimeng Zhang, Xin Chen, Jinghan Jia, Yihua Zhang 等NeurIPS 2024 · 被引用 200 次
