Selective Fine-Tuning for Targeted and Robust Concept Unlearning
Mansi Mansi, Avinash Kori, Francesca Toni, Soteris Demetriou
摘要
Text-guided diffusion models are used by millions of users, but can be easily exploited to produce harmful content. Concept unlearning methods have emerged as a principled model-level defense, but state-of-the-art methods remain brittle: they are vulnerable to adversarial prompts, degrade generation quality for benign concepts, and lack interpretability. Concept localization approaches improve on concept unlearning through more targeted and interpretable selective fine-tuning, but current techniques are still limited: they are static, which results in suboptimal utility and weak robustness. In order to tackle these challenges, we propose TRuST (Targeted Robust Selective fine-Tuning), a novel approach for dynamically localizing target concept neurons and unlearning them demonstrated with two new optimization objectives and a selective fine-tuning method. We evaluate TRuST under both static and adaptive adversaries, and show experimentally, against a number of state-of-the-art baselines, that TRuST achieves up to 14× reduction in Attack Success Rate (ASR) across both adversary types, preserves generation quality (ΔFID = 0.016, CLIP = 30.95), and converges in up to 22× fewer fine-tuning steps than the state-of-the-art. Our method achieves robust and interpretable unlearning of not only individual concepts but also combinations of concepts and conditional concepts, without any task-specific regularization. Content Warning: This paper contains examples of unsafe and explicit visual content and language included solely for research and evaluation purposes.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper37
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- Certified Data Removal from Machine Learning ModelsChuan Guo, Tom Goldstein, Awni Y. Hannun, Laurens van der MaatenICML 2020 · 被引用 633 次
相关 Paper
- ConceptPrune: Concept Editing in Diffusion Models via Skilled Neuron PruningRuchika Chavhan, Da Li, Timothy M. HospedalesICLR 2025 · 被引用 2 次
- Adaptive Median Smoothing: Adversarial Defense for Unlearned Text-to-Image Diffusion Models at Inference TimeXiaoxuan Han, Songlin Yang, Wei Wang, Yang Li 等ICML 2025
- Meta-Unlearning on Diffusion Models: Preventing Relearning Unlearned ConceptsHongcheng Gao, Tianyu Pang, Chao Du, Taihang Hu 等ICCV 2025 · 被引用 4 次
- Defensive Unlearning with Adversarial Training for Robust Concept Erasure in Diffusion ModelsYimeng Zhang, Xin Chen, Jinghan Jia, Yihua Zhang 等NeurIPS 2024 · 被引用 200 次
- Concept Replacer: Replacing Sensitive Concepts in Diffusion Models via Precision LocalizationLingyun Zhang, Yu Xie, Yanwei Fu, Ping ChenCVPR 2025
