VL-InterpreT: An Interactive Visualization Tool for Interpreting Vision-Language Transformers
Estelle Aflalo, Meng Du, Shao-Yen Tseng, Yongfei Liu, Chenfei Wu, Nan Duan, Vasudev Lal
Abstract
Breakthroughs in transformer-based models have revolutionized not only the NLP field, but also vision and multimodal systems. However, although visualization and interpretability tools have become available for NLP models, internal mechanisms of vision and multimodal transformers remain largely opaque. With the success of these transformers, it is increasingly critical to understand their inner workings, as unraveling these black-boxes will lead to more capable and trustworthy models. To contribute to this quest, we propose VL-InterpreT, which provides novel interactive visualizations for interpreting the attentions and hidden representations in multimodal transformers. VL-InterpreT is a task agnostic and integrated tool that (1) tracks a variety of statistics in attention heads throughout all layers for both vision and language components, (2) visualizes cross-modal and intra-modal attentions through easily readable heatmaps, and (3) plots the hidden representations of vision and language tokens as they pass through the transformer layers. In this paper, we demonstrate the functionalities of VL-InterpreT through the analysis of KD-VLP, an end-to-end pretraining vision-language multimodal transformer-based model, in the tasks of Visual Commonsense Reasoning (VCR) and WebQA, two visual question answering benchmarks. Furthermore, we also present a few interesting findings about multimodal transformer behaviors that were learned through our tool.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e75cda6f-16fd-4333-b3fc-70508d08f2fbCited by top-tier papers21
- AttentionViz: A Global View of Transformer AttentionCatherine Yeh, Yida Chen, Aoyu Wu, Cynthia Chen et al.IEEE VIS 2023 · 78 citations
- MultiViz: Towards Visualizing and Understanding Multimodal ModelsPaul Pu Liang, Yiwei Lyu, Gunjan Chhablani, Nihal Jain et al.ICLR 2023 · 15 citations
- Grounding Visual Illusions in Language: Do Vision-Language Models Perceive Illusions Like Humans?Yichi Zhang, Jiayi Pan, Yuchen Zhou, Rui Pan et al.EMNLP 2023 · 9 citations
- Dissecting Multimodality in VideoQA Transformer Models by Impairing Modality FusionIshaan Singh Rawal, Alexander Matyasko, Shantanu Jaiswal, Basura Fernando et al.ICML 2024 · 8 citations
- Faithful and Accurate Self-Attention Attribution for Message Passing Neural Networks via the Computation Tree ViewpointYong-Min Shin, Siqing Li, Xin Cao, Won-Yong ShinAAAI 2025 · 6 citations
Builds on10
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 2,196 citations
- On the Relationship between Self-Attention and Convolutional LayersJean-Baptiste Cordonnier, Andreas Loukas, Martin JaggiICLR 2020 · 629 citations
- Generic Attention-model Explainability for Interpreting Bi-Modal and Encoder-Decoder TransformersHila Chefer, Shir Gur, Lior WolfICCV 2021 · 451 citations
Related papers
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li et al.ICLR 2020 · 1,825 citations
- Do Vision-Language Transformers Exhibit Visual Commonsense? An Empirical Study of VCRZhenyang Li, Yangyang Guo, Kejie Wang, Xiaolin Chen et al.ACM MM 2023 · 11 citations
- Probing Inter-modality: Visual Parsing with Self-Attention for Vision-and-Language Pre-trainingHongwei Xue, Yupan Huang, Bei Liu, Houwen Peng et al.NeurIPS 2021 · 100 citations
- Understanding Language Prior of LVLMs by Contrasting Chain-of-EmbeddingLin Long, Changdae Oh, Seongheon Park, Sharon LiICLR 2026 · 14 citations
- Unified Vision-Language Pre-Training for Image Captioning and VQALuowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu et al.AAAI 2020 · 1,047 citations
