DSS: Dynamic Semantic Steering for Robust Concept Erasure in Diffusion Models
Qinghui Gong, Zhengchun Zhou, Hua Meng, Yihuai Liang, Yuxuan Zhang
摘要
Text-to-image (T2I) diffusion models have introduced new security risks, as adversaries can exploit flexible text prompts to induce the generation of sensitive or policy-violating content (e.g., NSFW or copyrighted concepts). Concept erasure has emerged as a promising defense, aiming to suppress targeted semantics while preserving benign generation. However, existing approaches face a fundamental trade-off: training-based methods are costly and inflexible to emerging threats, while inference-time interventions often rely on unconstrained feature manipulation, leading to over-correction, semantic drift, and degraded utility. More critically, they lack an explicit understanding of local semantic structure, resulting in unstable behavior under adversarial or compositional prompts. In this paper, we propose Dynamic Semantic Steering (DSS), a training-free, inference-time defense framework for robust and controllable concept erasure under prompt-level adversarial settings. DSS introduces a geometry-aware formulation that explicitly models local semantic neighborhoods and constrains feature updates within this structure, which is key to achieving precise suppression without sacrificing benign semantics. Specifically, DSS (i) automatically identifies benign semantic anchors via density-based boundary modeling, and (ii) performs context-aware, constrained feature correction using cross-attention signals with a closed-form solution. Extensive experiments demonstrate that DSS achieves strong and consistent suppression across diverse concept categories and adversarial prompts, reaching an average erasure rate of 91.0%, outperforming prior defenses (18.6%--85.9%). At the same time, DSS substantially reduces semantic drift and preserves generation fidelity.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion ModelsAlexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam 等ICML 2022 · 被引用 4,691 次
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann 等ICLR 2024 · 被引用 4,569 次
- DPM-Solver: A Fast ODE Solver for Diffusion Probabilistic Model Sampling in Around 10 StepsCheng Lu, Yuhao Zhou, Fan Bao, Jianfei Chen 等NeurIPS 2022 · 被引用 2,653 次
相关 Paper
- GenErase: Generalizable and Semantically-Aware Concept Erasure in Diffusion ModelsKorada Sri Vardhana, Soma BiswasCVPR 2026
- TRCE: Towards Reliable Malicious Concept Erasure in Text-to-Image Diffusion ModelsRuidong Chen, Honglin Guo, Lanjun Wang, Chenyu Zhang 等ICCV 2025 · 被引用 17 次
- MapRoute:Precise-Concept Erasing Mappers via Semantic RoutingSihao Li, Baixi Baixi, Shuohong Xia, Yunyun YangCVPR 2026
- GrOCE : Graph-Guided Online Concept Erasure for Text-to-Image Diffusion ModelsNing Han, Zhenyu Ge, Feng Han, Yuhua Sun 等CVPR 2026 · 被引用 3 次
- Semantic Surgery: Zero-Shot Concept Erasure in Diffusion ModelsLexiang Xiong, Chengyu Liu, Jingwen Ye, Yan Liu 等NeurIPS 2025 · 被引用 8 次
