CLIPErase: Efficient Unlearning of Visual-Textual Associations in CLIP
Tianyu Yang, Lisen Dai, Xiangqi Wang, Minhao Cheng, Yapeng Tian, Xiangliang Zhang
Abstract
Machine unlearning (MU) has gained significant attention as a means to remove the influence of specific data from a trained model without requiring full retraining. While progress has been made in unimodal domains like text and image classification, unlearning in multimodal models remains relatively underexplored. In this work, we address the unique challenges of unlearning in CLIP, a prominent multimodal model that aligns visual and textual representations. We introduce CLIPErase, a novel approach that disentangles and selectively forgets both visual and textual associations, ensuring that unlearning does not compromise model performance. CLIPErase consists of three key modules: a Forgetting Module that disrupts the associations in the forget set, a Retention Module that preserves performance on the retain set, and a Consistency Module that maintains consistency with the original model. Extensive experiments on CIFAR-100, Flickr30K, and Conceptual 12M across five CLIP downstream tasks, as well as an evaluation on diffusion models, demonstrate that CLIPErase effectively removes designated associations from multimodal samples in downstream tasks, while preserving the model's performance on the retain set after unlearning. The project's code is available at: https://tianyuyang-anna.github.io/ClipErase-ACL/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- SALMUBench: A Benchmark for Sensitive Association-Level Multimodal UnlearningCai Selvas-Sala, Lei Kang, Lluís GómezCVPR 2026 · 4 citations
- Towards Reasoning-Preserving Unlearning in Multimodal Large Language ModelsHongji Li, Manjiang Yu, Junchi Yao, PRIYANKA SINGH et al.CVPR 2026 · 3 citations
- Cross-Modal Unlearning via Influential Neuron Path Editing in Multimodal Large Language ModelsKunhao Li, Wenhao Li, Di Wu, Lei Yang et al.AAAI 2026 · 2 citations
- VL-Eraser: Vacuum Distillation for Machine Unlearning in Vision-Language ModelsYili Wang, Lu Dai, Tairan Huang, Yijie Xu et al.CVPR 2026
- POUR: A Provably Optimal Method for Unlearning Representation via Neural CollapseAnjie Le, Can Peng, Yuyuan Liu, Alison NobleCVPR 2026
Builds on15
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty et al.NeurIPS 2021 · 2,985 citations
- Erasing Concepts from Diffusion ModelsRohit Gandikota, Joanna Materzynska, Jaden Fiotto-Kaufman, David BauICCV 2023 · 536 citations
Related papers
- Captured by Captions: On Memorization and its Mitigation in CLIP ModelsWenhao Wang, Adam Dziedzic, Grace C. Kim, Michael Backes et al.ICLR 2025
- Causal Disentanglement and Cross-Modal Alignment for Enhanced Few-Shot LearningTianjiao Jiang, Zhen Zhang, Yuhang Liu, Javen Qinfeng ShiICCV 2025 · 3 citations
- Multimodal Dataset Distillation Made Simple by Prototype-Guided Data SynthesisJunhyeok Choi, Sangwoo Mo, Minwoo ChaeICLR 2026
- CLIPPO: Image-and-Language Understanding from Pixels OnlyMichael Tschannen, Basil Mustafa, Neil HoulsbyCVPR 2023
- SmartCLIP: Modular Vision-language Alignment with Identification GuaranteesShaoan Xie, Lingjing Kong, Yujia Zheng, Yu Yao et al.CVPR 2025
