DSS: Dynamic Semantic Steering for Robust Concept Erasure in Diffusion Models
Qinghui Gong, Zhengchun Zhou, Hua Meng, Yihuai Liang, Yuxuan Zhang
Abstract
Text-to-image (T2I) diffusion models have introduced new security risks, as adversaries can exploit flexible text prompts to induce the generation of sensitive or policy-violating content (e.g., NSFW or copyrighted concepts). Concept erasure has emerged as a promising defense, aiming to suppress targeted semantics while preserving benign generation. However, existing approaches face a fundamental trade-off: training-based methods are costly and inflexible to emerging threats, while inference-time interventions often rely on unconstrained feature manipulation, leading to over-correction, semantic drift, and degraded utility. More critically, they lack an explicit understanding of local semantic structure, resulting in unstable behavior under adversarial or compositional prompts. In this paper, we propose Dynamic Semantic Steering (DSS), a training-free, inference-time defense framework for robust and controllable concept erasure under prompt-level adversarial settings. DSS introduces a geometry-aware formulation that explicitly models local semantic neighborhoods and constrains feature updates within this structure, which is key to achieving precise suppression without sacrificing benign semantics. Specifically, DSS (i) automatically identifies benign semantic anchors via density-based boundary modeling, and (ii) performs context-aware, constrained feature correction using cross-attention signals with a closed-form solution. Extensive experiments demonstrate that DSS achieves strong and consistent suppression across diverse concept categories and adversarial prompts, reaching an average erasure rate of 91.0%, outperforming prior defenses (18.6%--85.9%). At the same time, DSS substantially reduces semantic drift and preserves generation fidelity.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion ModelsAlexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam et al.ICML 2022 · 4,691 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
- DPM-Solver: A Fast ODE Solver for Diffusion Probabilistic Model Sampling in Around 10 StepsCheng Lu, Yuhao Zhou, Fan Bao, Jianfei Chen et al.NeurIPS 2022 · 2,653 citations
Related papers
- GenErase: Generalizable and Semantically-Aware Concept Erasure in Diffusion ModelsKorada Sri Vardhana, Soma BiswasCVPR 2026
- TRCE: Towards Reliable Malicious Concept Erasure in Text-to-Image Diffusion ModelsRuidong Chen, Honglin Guo, Lanjun Wang, Chenyu Zhang et al.ICCV 2025 · 17 citations
- MapRoute:Precise-Concept Erasing Mappers via Semantic RoutingSihao Li, Baixi Baixi, Shuohong Xia, Yunyun YangCVPR 2026
- GrOCE : Graph-Guided Online Concept Erasure for Text-to-Image Diffusion ModelsNing Han, Zhenyu Ge, Feng Han, Yuhua Sun et al.CVPR 2026 · 3 citations
- Semantic Surgery: Zero-Shot Concept Erasure in Diffusion ModelsLexiang Xiong, Chengyu Liu, Jingwen Ye, Yan Liu et al.NeurIPS 2025 · 8 citations
