Modality-Aware Neuron Pruning for Unlearning in Multimodal Large Language Models
Zheyuan Liu, Guangyao Dou, Xiangchi Yuan, Chunhui Zhang, Zhaoxuan Tan, Meng Jiang
Abstract
Generative models such as Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs) trained on massive datasets can lead them to memorize and inadvertently reveal sensitive information, raising ethical and privacy concerns. While some prior works have explored this issue in the context of LLMs, it presents a unique challenge for MLLMs due to the entangled nature of knowledge across modalities, making comprehensive unlearning more difficult. To address this challenge, we propose Modality Aware Neuron Unlearning (MANU), a novel unlearning framework for MLLMs designed to selectively clip neurons based on their relative importance to the targeted forget data, curated for different modalities. Specifically, MANU consists of two stages: important neuron selection and selective pruning. The first stage identifies and collects the most influential neurons across modalities relative to the targeted forget knowledge, while the second stage is dedicated to pruning those selected neurons. MANU effectively isolates and removes the neurons that contribute most to the forget data within each modality, while preserving the integrity of retained knowledge. Our experiments conducted across various MLLM architectures illustrate that MANU can achieve a more balanced and comprehensive unlearning in each modality without largely affecting the overall model utility. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 85ff1c7c-35d3-410d-b710-17ec6b0af7faCited by top-tier papers9
- Superficial Self-Improved Reasoners Benefit from Model MergingXiangchi Yuan, Chunhui Zhang, Zheyuan Liu, Dachuan Shi et al.EMNLP 2025 · 15 citations
- Machine Unlearning via Task Simplex ArithmeticJunhao Dong, Hao Zhu, Yifei Zhang, Xinghua Qu et al.NeurIPS 2025 · 10 citations
- Model Unlearning via Sparse Autoencoder Subspace Guided ProjectionsXu Wang, Zihao Li, Benyou Wang, Yan Hu et al.EMNLP 2025 · 9 citations
- Behavior Knowledge Merge in Reinforced Agentic ModelsXiangchi Yuan, Dachuan Shi, Chunhui Zhang, Zheyuan Liu et al.ACL 2026 · 7 citations
- Growing Through Experience: Scaling Episodic Grounding in Language ModelsChunhui Zhang, Sirui Wang, Zhongyu Ouyang, Xiangchi Yuan et al.ACL 2025 · 6 citations
Builds on19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski et al.USENIX Security 2021 · 2,866 citations
Related papers
- Cross-Modal Unlearning via Influential Neuron Path Editing in Multimodal Large Language ModelsKunhao Li, Wenhao Li, Di Wu, Lei Yang et al.AAAI 2026 · 2 citations
- Knowledge Externalization: Reversible Unlearning and Modular Retrieval in Multimodal Large Language ModelsJiaqi Li, Zihan You, Ruoyan Shen, Shenyu Zhang et al.ICLR 2026
- ASRU: Activation Steering Meets Reinforcement Unlearning for Multimodal Large Language ModelsJiahui Guang, Haiyan Wang, Yingjie Zhu, Cuiyun Gao et al.ICML 2026
- LOTUS: Evolving Multimodal Unlearning via Hyperbolic Entailment and Lorentz TransportZekun Wang, Jingjie Zeng, Yingxu Li, Hongfei Lin et al.ACL 2026
- SUA: Stealthy Multimodal Large Language Model Unlearning AttackXianren Zhang, Hui Liu, Delvin Ce Zhang, Xianfeng Tang et al.EMNLP 2025
