Selective Fine-Tuning for Targeted and Robust Concept Unlearning
Mansi Mansi, Avinash Kori, Francesca Toni, Soteris Demetriou
Abstract
Text-guided diffusion models are used by millions of users, but can be easily exploited to produce harmful content. Concept unlearning methods have emerged as a principled model-level defense, but state-of-the-art methods remain brittle: they are vulnerable to adversarial prompts, degrade generation quality for benign concepts, and lack interpretability. Concept localization approaches improve on concept unlearning through more targeted and interpretable selective fine-tuning, but current techniques are still limited: they are static, which results in suboptimal utility and weak robustness. In order to tackle these challenges, we propose TRuST (Targeted Robust Selective fine-Tuning), a novel approach for dynamically localizing target concept neurons and unlearning them demonstrated with two new optimization objectives and a selective fine-tuning method. We evaluate TRuST under both static and adaptive adversaries, and show experimentally, against a number of state-of-the-art baselines, that TRuST achieves up to 14× reduction in Attack Success Rate (ASR) across both adversary types, preserves generation quality (ΔFID = 0.016, CLIP = 30.95), and converges in up to 22× fewer fine-tuning steps than the state-of-the-art. Our method achieves robust and interpretable unlearning of not only individual concepts but also combinations of concepts and conditional concepts, without any task-specific regularization. Content Warning: This paper contains examples of unsafe and explicit visual content and language included solely for research and evaluation purposes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 579cdcb3-92c1-4a38-86dc-c1136427ae58Builds on37
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Certified Data Removal from Machine Learning ModelsChuan Guo, Tom Goldstein, Awni Y. Hannun, Laurens van der MaatenICML 2020 · 633 citations
Related papers
- ConceptPrune: Concept Editing in Diffusion Models via Skilled Neuron PruningRuchika Chavhan, Da Li, Timothy M. HospedalesICLR 2025 · 2 citations
- Adaptive Median Smoothing: Adversarial Defense for Unlearned Text-to-Image Diffusion Models at Inference TimeXiaoxuan Han, Songlin Yang, Wei Wang, Yang Li et al.ICML 2025
- Meta-Unlearning on Diffusion Models: Preventing Relearning Unlearned ConceptsHongcheng Gao, Tianyu Pang, Chao Du, Taihang Hu et al.ICCV 2025 · 4 citations
- Defensive Unlearning with Adversarial Training for Robust Concept Erasure in Diffusion ModelsYimeng Zhang, Xin Chen, Jinghan Jia, Yihua Zhang et al.NeurIPS 2024 · 200 citations
- Concept Replacer: Replacing Sensitive Concepts in Diffusion Models via Precision LocalizationLingyun Zhang, Yu Xie, Yanwei Fu, Ping ChenCVPR 2025
