Peeling Back the Layers: Interpreting the Storytelling of ViT
Jingjie Zeng, Zhihao Yang, Qi Yang, Liang Yang, Hongfei Lin
Abstract
By integrating various modules with the Visual Transformer (ViT), we facilitate a interpretation of image processing across each layer and attention head. This method allows us to explore the connections both within and across the layers, enabling a analysis of how images are processed at different layers. Conducting a analysis of the contributions from each layer and attention head, shedding light on the intricate interactions and functionalities within the model's layers. This in-depth exploration not only highlights the visual cues between layers but also examines their capacity to navigate the transition from abstract concepts to tangible objects. It unveils the model's mechanism to building an understanding of images, providing a strategy for adjusting attention heads between layers, thus enabling targeted pruning and enhancement of performance for specific tasks. Our research indicates that achieving a scalable understanding of transformer models is within reach, offering ways for the refinement and enhancement of such models.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers1
Ask how each one uses itRelated papers
- Decomposing and Interpreting Image Representations via Text in ViTs Beyond CLIPSriram Balasubramanian, Samyadeep Basu, Soheil FeiziNeurIPS 2024 · 26 citations
- Attention Guided CAM: Visual Explanations of Vision Transformer Guided by Self-AttentionSaebom Leem, Hyunseok SeoAAAI 2024 · 40 citations
- Revisiting Vision Transformer from the View of Path EnsembleShuning Chang, Pichao Wang, Hao Luo, Fan Wang et al.ICCV 2023 · 8 citations
- Dissecting Query-Key Interaction in Vision TransformersXu Pan, Aaron Philip, Ziqian Xie, Odelia SchwartzNeurIPS 2024 · 18 citations
- Multimodal Language Models See Better When They Look ShallowerHaoran Chen, Junyan Lin, Xinghao Chen, Yue Fan et al.EMNLP 2025
