Large Multi-modal Models Can Interpret Features in Large Multi-modal Models
Kaichen Zhang, Yifei Shen, Bo Li, Ziwei Liu
Abstract
Recent advances in Large Multimodal Models (LMMs) lead to significant breakthroughs in both academia and industry. One question that arises is how we, as humans, can understand their internal neural representations. This paper takes an initial step towards addressing this question by presenting a versatile framework to identify and interpret the semantics within LMMs. Specifically, 1) we first apply a Sparse Autoencoder(SAE) to disentangle the representations into human understandable features. 2) We then present an automatic interpretation framework to interpreted the open-semantic features learned in SAE by the LMMs themselves. We employ this framework to analyze the LLaVA-NeXT-8B model using the LLaVA-OV-72B model, demonstrating that these features can effectively steer the model's behavior. Our results contribute to a deeper understanding of why LMMs excel in specific tasks, including EQ tests, and illuminate the nature of their mistakes along with potential strategies for their rectification. These findings offer new insights into the internal mechanisms of LMMs and suggest parallels with the cognitive processes of the human brain.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e64d39bb-fd6c-42f5-9fd9-4129b7fd3fe8Cited by top-tier papers11
- Sparse Autoencoders Learn Monosemantic Features in Vision-Language ModelsMateusz Pach, Shyamgopal Karthik, Quentin Bouniot, Serge J. Belongie et al.NeurIPS 2025 · 79 citations
- SUB: Benchmarking CBM Generalization via Synthetic Attribute SubstitutionsJessica Bader, Leander Girrbach, Stephan Alaniz, Zeynep AkataICCV 2025 · 8 citations
- Improving Sparse Autoencoder with Dynamic AttentionDongsheng Wang, Jinsen Zhang, Dawei Su, Hui HuangCVPR 2026 · 2 citations
- Language Models Can Explain Visual Features via SteeringJavier Ferrando, Enrique Lopez-Cuena, Pablo Agustin Martin-Torres, Daniel Hinjos et al.CVPR 2026 · 2 citations
- Dictionary-Aligned Concept Control for Safeguarding Multimodal LLMsJinqi Luo, Jinyu Yang, Tal Neiman, Lei Fan et al.CVPR 2026 · 1 citation
Builds on11
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
- SGLang: Efficient Execution of Structured Language Model ProgramsLianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun et al.NeurIPS 2024 · 1,586 citations
Related papers
- ConceptViz: A Visual Analytics Approach for Exploring Concepts in Large Language ModelsHaoxuan Li, Zhen Wen, Qiqi Jiang, Chenxiao Li et al.IEEE VIS 2025 · 3 citations
- SAE-V: Interpreting Multimodal Models for Enhanced AlignmentHantao Lou, Changye Li, Jiaming Ji, Yaodong YangICML 2025
- A Concept-Based Explainability Framework for Large Multimodal ModelsJayneel Parekh, Pegah Khayatan, Mustafa Shukor, Alasdair Newson et al.NeurIPS 2024 · 48 citations
- Neuro-Fuzzy Concept Learning for Interpretable Large Multimodal ModelsRitik Mishra, Vanshika Gupta, M. Sajid, M. TanveerICML 2026
- Enhancing Few-Shot Vision-Language Classification With Large Multimodal Model FeaturesChancharik Mitra, Brandon Huang, Tianning Chai, Zhiqiu Lin et al.ICCV 2025 · 2 citations
