Targeted Forgetting of Image Subgroups in CLIP Models
Zeliang Zhang, Gaowen Liu, Charles Fleming, Ramana Rao Kompella, Chenliang Xu
Abstract
Foundation models (FMs) such as CLIP have demonstrated impressive zero-shot performance across various tasks by leveraging large-scale, unsupervised pre-training. However, they often inherit harmful or unwanted knowledge from noisy internet-sourced datasets, compromising their reliability in real-world applications. Existing model unlearning methods either rely on access to pre-trained datasets or focus on coarse-grained unlearning (e.g., entire classes), leaving a critical gap for fine-grained unlearning. In this paper, we address the challenging scenario of selectively forgetting specific portions of knowledge within a class-without access to pre-trained data-while preserving the model's overall performance. We propose a novel three-stage approach that progressively unlearns targeted knowledge while mitigating over-forgetting. It consists of (1) a forgetting stage to fine-tune the CLIP on samples to be forgotten, (2) a reminding stage to restore performance on retained samples, and (3) a restoring stage to recover zero-shot capabilities using model souping. Additionally, we introduce knowledge distillation to handle the distribution disparity between forgetting/retaining samples and unseen pre-trained data. Extensive experiments on CIFAR-10, ImageNet-1K, and style datasets demonstrate that our approach effectively unlearns specific subgroups while maintaining strong zeroshot performance on semantically similar subgroups and other categories, significantly outperforming baseline unlearning methods, which lose effectiveness under the CLIP unlearning setting.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cb45c032-d6d8-4a0c-ad7d-343d7f1061b1Cited by top-tier papers2
- SafeR-CLIP: Mitigating NSFW Content in Vision-Language Models While Preserving Pre-Trained KnowledgeAdeel Yousaf, Joseph Fioresi, James Beetham, Amrit Singh Bedi et al.AAAI 2026 · 1 citation
- Selective Amnesia using Contrastive Subnet Erasure for Class Level Unlearning in Vision ModelsVishal Pramanik, Maisha Maliha, Susmit Jha, Alvaro Velasquez et al.CVPR 2026
Builds on26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs et al.ICML 2022 · 1,464 citations
- Machine UnlearningLucas Bourtoule, Varun Chandrasekaran, Christopher A. Choquette-Choo, Hengrui Jia et al.S&P 2021 · 1,381 citations
- Remember What You Want to Forget: Algorithms for Machine UnlearningAyush Sekhari, Jayadev Acharya, Gautam Kamath, Ananda Theertha SureshNeurIPS 2021 · 516 citations
- Model Sparsity Can Simplify Machine UnlearningJinghan Jia, Jiancheng Liu, Parikshit Ram, Yuguang Yao et al.NeurIPS 2023 · 293 citations
Related papers
- Decoupled Distillation to Erase: A General Unlearning Method for Any Class-centric TasksYu Zhou, Dian Zheng, Qijie Mo, Renjie Lu et al.CVPR 2025
- Rethinking Few Shot CLIP Benchmarks: A Critical Analysis in the Inductive SettingAlexey Kravets, Da Chen, Vinay P. NamboodiriICCV 2025 · 1 citation
- Black-Box ForgettingYusuke Kuwana, Yuta Goto, Takashi Shibata, Go IrieNeurIPS 2024 · 6 citations
- Unlearning’s Blind Spots: Over‑Unlearning and Prototypical Relearning AttackSeungBum Ha, Saerom Park, Sung Whan YoonICML 2026 · 2 citations
- Forget What Has Seen: Selective Concept Unlearning in Segmentation Foundation ModelsMiaozeng Du, Jiaqi Li, Sirui Pan, Yi Zhan et al.AAAI 2026
