Towards Neuron Attributions in Multi-Modal Large Language Models
Junfeng Fang, Zac Bi, Ruipeng Wang, Houcheng Jiang, Yuan Gao, Kun Wang, An Zhang, Jie Shi, Xiang Wang, Tat-Seng Chua
Abstract
As Large Language Models (LLMs) demonstrate impressive capabilities, demys-tifying their internal mechanisms becomes increasingly vital. Neuron attribution, which attributes LLM outputs to specific neurons to reveal the semantic properties they learn, has emerged as a key interpretability approach. However, while neuron attribution has made significant progress in deciphering text-only LLMs, its application to Multimodal LLMs (MLLMs) remains less explored. To address this gap, we propose a novel N euron A ttribution method tailored for M LLMs, termed NAM . Specifically, NAM not only reveals the modality-specific semantic knowledge learned by neurons within MLLMs, but also highlights several intriguing properties of neurons, such as cross-modal invariance and semantic sensitivity. These properties collectively elucidate the inner workings mechanism of MLLMs, providing a deeper understanding of how MLLMs process and generate multi-modal content. Through theoretical analysis and empirical validation, we demonstrate the efficacy of NAM and the valuable insights it offers. Furthermore, leveraging NAM, we introduce a multi-modal knowledge editing paradigm, underscoring the practical significance of our approach for downstream applications of MLLMs. Our code is available at https://github.com/littlelittlenine/NAM_1.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 165bcd4f-cb51-4974-9f5a-9dc6c8e4ceedCited by top-tier papers9
- β-DPO: Direct Preference Optimization with Dynamic βJunkang Wu, Yuexiang Xie, Zhengyi Yang, Jiancan Wu et al.NeurIPS 2024 · 114 citations
- Discovering and Causally Validating Emotion-Sensitive Neurons in Large Audio-Language ModelsXiutian Zhao, Björn W. Schuller, Berrak SismanACL 2026 · 5 citations
- From Concepts to Components: Concept-Agnostic Attention Module Discovery in TransformersJingtong Su, Julia Kempe, Karen UllrichICLR 2026 · 4 citations
- Reliable Lifelong Multimodal Editing: Conflict-Aware Retrieval Meets Multi-Level GuidanceQiang Zhang, Fanrui Zhang, Jiawei Liu, Ming Hu et al.NeurIPS 2025 · 1 citation
- Deciphering Functions of Neurons in Vision-Language ModelsJiaqi Xu, Cuiling Lan, Yan LuACM MM 2025
Builds on29
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
Related papers
- Cross-Modal Unlearning via Influential Neuron Path Editing in Multimodal Large Language ModelsKunhao Li, Wenhao Li, Di Wu, Lei Yang et al.AAAI 2026 · 2 citations
- Can Knowledge be Transferred from Unimodal to Multimodal? Investigating the Transitivity of Multimodal Knowledge EditingLingyong Fang, Xinzhong Wang, Depeng Wang, Zongru Wu et al.ICCV 2025 · 4 citations
- Correct When Paired, Wrong When Split: Decoupling and Editing Modality-Specific Neurons in MLLMsTingchao Fu, Wenkai Wang, Fanxiao Li, Huadong Zhang et al.ACL 2026
- NeuCon-ICE: Neuron-Level Controllable In-Context Editing for Multimodal Large Language ModelsChao Jiang, Jinzhi Liao, Xiang ZhaoSIGIR 2026
- Can We Debias Multimodal Large Language Models via Model Editing?Zecheng Wang, Xinye Li, Zhanyue Qin, Chunshan Li et al.ACM MM 2024 · 2 citations
