Neuro-Fuzzy Concept Learning for Interpretable Large Multimodal Models
Ritik Mishra, Vanshika Gupta, M. Sajid, M. Tanveer
摘要
Large Multimodal Models (LMMs) integrate unimodal encoders with Large Language Models (LLMs) to execute complex multimodal tasks. Despite progress in the field, understanding the internal representations of these models through interpretable logic remains an open problem. To address this, we present a framework utilizing a Human-Inspired (Neuro-fuzzy) approach for learning token representations. In this method, we leverage fuzzy rules to compute activation firing strengths, which are subsequently defuzzified to extract distinct concepts. This mechanism allows for the interpretation of learned representations directly through explicit logic. Consequently, we derive "multimodal concepts" that are both semantically coherent and interpretable. We validate our approach through rigorous qualitative and quantitative experiments, demonstrating the utility of these concepts in interpreting test samples. Additionally, we evaluate the disentanglement of the learned concepts and the efficacy of their grounding in both visual and textual domains. The source code is available at https://github.com/ mtanveer1/Neuro-FeX-LMM .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper22
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
相关 Paper
- A Concept-Based Explainability Framework for Large Multimodal ModelsJayneel Parekh, Pegah Khayatan, Mustafa Shukor, Alasdair Newson 等NeurIPS 2024 · 被引用 48 次
- Large Multi-modal Models Can Interpret Features in Large Multi-modal ModelsKaichen Zhang, Yifei Shen, Bo Li, Ziwei LiuICCV 2025 · 被引用 2 次
- Symbol-LLM: Leverage Language Models for Symbolic System in Visual Human Activity ReasoningXiaoqian Wu, Yonglu Li, Jianhua Sun, Cewu LuNeurIPS 2023 · 被引用 40 次
- Towards Neuron Attributions in Multi-Modal Large Language ModelsJunfeng Fang, Zac Bi, Ruipeng Wang, Houcheng Jiang 等NeurIPS 2024 · 被引用 16 次
- BrainFLORA: Uncovering Brain Concept Representation via Multimodal Neural EmbeddingsDongyang Li, Haoyang Qin, Mingyang Wu, Chen Wei 等ACM MM 2025 · 被引用 1 次
