Targeted Unlearning with Single Layer Unlearning Gradient
Zikui Cai, Yaoteng Tan, M. Salman Asif
摘要
The unauthorized generation of privacy-related and copyright-infringing content using generative-AI is becoming a significant concern for society, raising ethical, legal, and privacy issues that demand urgent attention. Recently, machine unlearning techniques have arisen that attempt to eliminate the influence of sensitive content used during model training, but they often require extensive updates in the model, reduce the utility of the models for unrelated content, and/or incur substantial computational costs. In this work, we propose a novel and efficient method called Single Layer Unlearning Gradient (SLUG), that can unlearn targeted information by updating a single targeted layer of a model using a one-time gradient computation. We introduce two metrics: layer importance and gradient alignment, to identify the appropriate layers for unlearning targeted information. Our method is highly modular and enables selective removal of multiple concepts from the generated outputs of widely used foundation models (e.g., CLIP), generative models (e.g., Stable Diffusion) and Vision-Language models. Our method shows effectiveness on a broad spectrum of concepts ranging from concrete (e.g., celebrity name, intellectual property figure, and object) to abstract (e.g., novel concept and artistic style). Our code is available at https://github.com/CSIPlab/SLUG .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- SAEmnesia: Erasing Concepts in Diffusion Models with Supervised Sparse AutoencodersEnrico Cassano, Riccardo Renzulli, Marco Nurisso, Mirko Zaffaroni 等ICML 2026 · 被引用 7 次
- Selective Fine-Tuning for Targeted and Robust Concept UnlearningMansi Mansi, Avinash Kori, Francesca Toni, Soteris DemetriouCCS 2026 · 被引用 2 次
- Beyond Sample-Level Forgetting: Improving Reliability in Multimodal UnlearningJianzhou Wang, Yirui Wu, Lixin Yuan, WENXIAO ZHANG 等ICML 2026
- Erased but Not Forgotten: How Backdoors Compromise Concept ErasureTobias Braun, Jonas Henry Grebe, Patrick Mohr Gordillo, Marcus Rohrbach 等ICML 2026
它引用的顶会 Paper13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Erasing Concepts from Diffusion ModelsRohit Gandikota, Joanna Materzynska, Jaden Fiotto-Kaufman, David BauICCV 2023 · 被引用 536 次
- SalUn: Empowering Machine Unlearning via Gradient-based Weight Saliency in Both Image Classification and GenerationChongyu Fan, Jiancheng Liu, Yihua Zhang, Eric Wong 等ICLR 2024 · 被引用 351 次
- Ablating Concepts in Text-to-Image Diffusion ModelsNupur Kumari, Bingliang Zhang, Sheng-Yu Wang, Eli Shechtman 等ICCV 2023 · 被引用 327 次
- Fast Machine Unlearning without Retraining through Selective Synaptic DampeningJack Foster, Stefan Schoepf, Alexandra BrintrupAAAI 2024 · 被引用 208 次
相关 Paper
- Boosting Alignment for Post-Unlearning Text-to-Image Generative ModelsMyeongseob Ko, Henry Li, Zhun Wang, Jonathan Patsenker 等NeurIPS 2024 · 被引用 22 次
- Image Can Bring Your Memory Back: A Novel Multi-Modal Guided Attack against Image Generation Model UnlearningRenyang Liu, Guanlin Li, Tianwei Zhang, See-Kiong NgICLR 2026 · 被引用 9 次
- The Illusion of Unlearning: The Unstable Nature of Machine Unlearning in Text-to-Image Diffusion ModelsNaveen George, Karthik Nandan Dasaraju, Rutheesh Reddy Chittepu, Konda Reddy MopuriCVPR 2025
- Holistic Unlearning Benchmark: A Multi-Faceted Evaluation for Text-to-Image Diffusion Model UnlearningSaemi Moon, Minjong Lee, Sangdon Park, Dongwoo KimICCV 2025
- Meta-Unlearning on Diffusion Models: Preventing Relearning Unlearned ConceptsHongcheng Gao, Tianyu Pang, Chao Du, Taihang Hu 等ICCV 2025 · 被引用 4 次
