Towards Neuron Attributions in Multi-Modal Large Language Models
Junfeng Fang, Zac Bi, Ruipeng Wang, Houcheng Jiang, Yuan Gao, Kun Wang, An Zhang, Jie Shi, Xiang Wang, Tat-Seng Chua
摘要
As Large Language Models (LLMs) demonstrate impressive capabilities, demys-tifying their internal mechanisms becomes increasingly vital. Neuron attribution, which attributes LLM outputs to specific neurons to reveal the semantic properties they learn, has emerged as a key interpretability approach. However, while neuron attribution has made significant progress in deciphering text-only LLMs, its application to Multimodal LLMs (MLLMs) remains less explored. To address this gap, we propose a novel N euron A ttribution method tailored for M LLMs, termed NAM . Specifically, NAM not only reveals the modality-specific semantic knowledge learned by neurons within MLLMs, but also highlights several intriguing properties of neurons, such as cross-modal invariance and semantic sensitivity. These properties collectively elucidate the inner workings mechanism of MLLMs, providing a deeper understanding of how MLLMs process and generate multi-modal content. Through theoretical analysis and empirical validation, we demonstrate the efficacy of NAM and the valuable insights it offers. Furthermore, leveraging NAM, we introduce a multi-modal knowledge editing paradigm, underscoring the practical significance of our approach for downstream applications of MLLMs. Our code is available at https://github.com/littlelittlenine/NAM_1.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- β-DPO: Direct Preference Optimization with Dynamic βJunkang Wu, Yuexiang Xie, Zhengyi Yang, Jiancan Wu 等NeurIPS 2024 · 被引用 114 次
- Discovering and Causally Validating Emotion-Sensitive Neurons in Large Audio-Language ModelsXiutian Zhao, Björn W. Schuller, Berrak SismanACL 2026 · 被引用 5 次
- From Concepts to Components: Concept-Agnostic Attention Module Discovery in TransformersJingtong Su, Julia Kempe, Karen UllrichICLR 2026 · 被引用 4 次
- Reliable Lifelong Multimodal Editing: Conflict-Aware Retrieval Meets Multi-Level GuidanceQiang Zhang, Fanrui Zhang, Jiawei Liu, Ming Hu 等NeurIPS 2025 · 被引用 1 次
- Deciphering Functions of Neurons in Vision-Language ModelsJiaqi Xu, Cuiling Lan, Yan LuACM MM 2025
它引用的顶会 Paper29
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
相关 Paper
- Cross-Modal Unlearning via Influential Neuron Path Editing in Multimodal Large Language ModelsKunhao Li, Wenhao Li, Di Wu, Lei Yang 等AAAI 2026 · 被引用 2 次
- Can Knowledge be Transferred from Unimodal to Multimodal? Investigating the Transitivity of Multimodal Knowledge EditingLingyong Fang, Xinzhong Wang, Depeng Wang, Zongru Wu 等ICCV 2025 · 被引用 4 次
- Correct When Paired, Wrong When Split: Decoupling and Editing Modality-Specific Neurons in MLLMsTingchao Fu, Wenkai Wang, Fanxiao Li, Huadong Zhang 等ACL 2026
- NeuCon-ICE: Neuron-Level Controllable In-Context Editing for Multimodal Large Language ModelsChao Jiang, Jinzhi Liao, Xiang ZhaoSIGIR 2026
- Can We Debias Multimodal Large Language Models via Model Editing?Zecheng Wang, Xinye Li, Zhanyue Qin, Chunshan Li 等ACM MM 2024 · 被引用 2 次
