Token Activation Map to Visually Explain Multimodal LLMs
Yi Li, Hualiang Wang, Xinpeng Ding, Haonan Wang, Xiaomeng Li
Abstract
Multimodal large language models (MLLMs) are broadly empowering various fields. Despite their advancements, the explainability of MLLMs remains less explored, hindering deeper understanding, model credibility, and effective visualization. Unlike conventional vision models (e.g., CNNs, ViTs, CLIP) that produce a single output, MLLMs generate sequences of tokens progressively, where each generated token depends on the previous context. Therefore, earlier context tokens can introduce redundant activations that interfere with the explanation of later tokens beyond their original information. Existing studies often overlook this issue, but our observations reveal that these redundant correlations can significantly hurt the reliability of explanations. To address this, we propose an estimated causal inference method to mitigate the interference of context to achieve high-quality MLLM explanation, with a novel rank Gaussian filter to further reduce activation noises. We term this method Token Activation Map (TAM) to highlight the consideration of interactions between tokens. TAM also indicates that it excels at explaining multiple tokens of MLLM, which is different from the Class Activation Map (CAM) for a single prediction. Our TAM method significantly outperforms existing SoTA methods, showcasing high-quality visualization results that can be utilized for various scenarios, such as object localization, failure case analysis, video visualization, MLLMs visual comparison, and model understanding (e.g., color, shape, action, location, visual reasoning, multi-turn conversation, etc). The code is available at github.com/xmed-lab/TAM.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers10
- Patch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMsYongyi Su, Haojie Zhang, Shijie Li, Nanqing Liu et al.ICLR 2026 · 22 citations
- Where MLLMs Attend and What They Rely On: Explaining Autoregressive Token GenerationRuoyu Chen, Xiaoqing Guo, Kangwei Liu, Siyuan Liang et al.CVPR 2026 · 19 citations
- Reading Images Like Texts: Sequential Image Understanding in Vision-Language ModelsYueyan Li, Chenggong Zhao, Zeyuan Zang, Caixia Yuan et al.ICLR 2026 · 3 citations
- Decomposed Attention Fusion in MLLMs for Training-free Video Reasoning SegmentationSu Ho Han, Jeongseok Hyun, Pilhyeon Lee, Minho Shim et al.ICLR 2026 · 2 citations
- One Token, Two Fates: A Unified Framework via Vision Token Manipulation Against MLLMs HallucinationZhan Fa, Yue Duan, Jian Zhang, Lei Qi et al.CVPR 2026 · 2 citations
Builds on18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
Related papers
- Diffusion-CAM: Faithful Visual Explanations for dMLLMsHaomin Zuo, Yidi Li, Luoxiao Yang, Xiaofeng ZhangACL 2026
- Task-Related Token Compression in Multimodal Large Language Models from an Explainability PerspectiveLei Lei, Jie Gu, Xiaokang Ma, Chu Tang et al.ICLR 2026 · 3 citations
- When MLLMs Meet Compression Distortion: A Coding Paradigm Tailored to MLLMsJinming Liu, Zhaoyang Jia, Jiahao Li, Bin Li et al.ICLR 2026 · 5 citations
- Exploring How Generative MLLMs Perceive More Than CLIP with the Same Vision EncoderSiting Li, Pang Wei Koh, Simon Shaolei DuACL 2025
- LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMsHaoran Lou, Chunxiao Fan, Ziyan Liu, Yuexin Wu et al.ICCV 2025 · 1 citation
