Targeted Forgetting of Image Subgroups in CLIP Models
Zeliang Zhang, Gaowen Liu, Charles Fleming, Ramana Rao Kompella, Chenliang Xu
摘要
Foundation models (FMs) such as CLIP have demonstrated impressive zero-shot performance across various tasks by leveraging large-scale, unsupervised pre-training. However, they often inherit harmful or unwanted knowledge from noisy internet-sourced datasets, compromising their reliability in real-world applications. Existing model unlearning methods either rely on access to pre-trained datasets or focus on coarse-grained unlearning (e.g., entire classes), leaving a critical gap for fine-grained unlearning. In this paper, we address the challenging scenario of selectively forgetting specific portions of knowledge within a class-without access to pre-trained data-while preserving the model's overall performance. We propose a novel three-stage approach that progressively unlearns targeted knowledge while mitigating over-forgetting. It consists of (1) a forgetting stage to fine-tune the CLIP on samples to be forgotten, (2) a reminding stage to restore performance on retained samples, and (3) a restoring stage to recover zero-shot capabilities using model souping. Additionally, we introduce knowledge distillation to handle the distribution disparity between forgetting/retaining samples and unseen pre-trained data. Extensive experiments on CIFAR-10, ImageNet-1K, and style datasets demonstrate that our approach effectively unlearns specific subgroups while maintaining strong zeroshot performance on semantically similar subgroups and other categories, significantly outperforming baseline unlearning methods, which lose effectiveness under the CLIP unlearning setting.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- SafeR-CLIP: Mitigating NSFW Content in Vision-Language Models While Preserving Pre-Trained KnowledgeAdeel Yousaf, Joseph Fioresi, James Beetham, Amrit Singh Bedi 等AAAI 2026 · 被引用 1 次
- Selective Amnesia using Contrastive Subnet Erasure for Class Level Unlearning in Vision ModelsVishal Pramanik, Maisha Maliha, Susmit Jha, Alvaro Velasquez 等CVPR 2026
它引用的顶会 Paper26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs 等ICML 2022 · 被引用 1,464 次
- Machine UnlearningLucas Bourtoule, Varun Chandrasekaran, Christopher A. Choquette-Choo, Hengrui Jia 等S&P 2021 · 被引用 1,381 次
- Remember What You Want to Forget: Algorithms for Machine UnlearningAyush Sekhari, Jayadev Acharya, Gautam Kamath, Ananda Theertha SureshNeurIPS 2021 · 被引用 516 次
- Model Sparsity Can Simplify Machine UnlearningJinghan Jia, Jiancheng Liu, Parikshit Ram, Yuguang Yao 等NeurIPS 2023 · 被引用 293 次
相关 Paper
- Decoupled Distillation to Erase: A General Unlearning Method for Any Class-centric TasksYu Zhou, Dian Zheng, Qijie Mo, Renjie Lu 等CVPR 2025
- Rethinking Few Shot CLIP Benchmarks: A Critical Analysis in the Inductive SettingAlexey Kravets, Da Chen, Vinay P. NamboodiriICCV 2025 · 被引用 1 次
- Black-Box ForgettingYusuke Kuwana, Yuta Goto, Takashi Shibata, Go IrieNeurIPS 2024 · 被引用 6 次
- Unlearning’s Blind Spots: Over‑Unlearning and Prototypical Relearning AttackSeungBum Ha, Saerom Park, Sung Whan YoonICML 2026 · 被引用 2 次
- Forget What Has Seen: Selective Concept Unlearning in Segmentation Foundation ModelsMiaozeng Du, Jiaqi Li, Sirui Pan, Yi Zhan 等AAAI 2026
