Generic Attention-model Explainability for Interpreting Bi-Modal and Encoder-Decoder Transformers
Hila Chefer, Shir Gur, Lior Wolf
Abstract
Transformers are increasingly dominating multi-modal reasoning tasks, such as visual question answering, achieving state-of-the-art results thanks to their ability to contextualize information using the self-attention and co-attention mechanisms. These attention modules also play a role in other computer vision tasks including object detection and image segmentation. Unlike Transformers that only use self-attention, Transformers with co-attention require to consider multiple attention maps in parallel in order to highlight the information that is relevant to the prediction in the model’s input. In this work, we propose the first method to explain prediction by any Transformer-based architecture, including bi-modal Transformers and Transformers with co-attentions. We provide generic solutions and apply these to the three most commonly used of these architectures: (i) pure self-attention, (ii) self-attention combined with co-attention, and (iii) encoder-decoder attention. We show that our method is superior to all existing methods which are adapted from single modality explainability. Our code is available at: https://github.com/hila-chefer/Transformer-MM-Explainability.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 93daef81-aecd-4979-b768-d79386a6c995Cited by top-tier papers137
- LERF: Language Embedded Radiance FieldsJustin Kerr, Chung Min Kim, Ken Goldberg, Angjoo Kanazawa et al.ICCV 2023 · 620 citations
- Attend-and-Excite: Attention-Based Semantic Guidance for Text-to-Image Diffusion ModelsHila Chefer, Yuval Alaluf, Yael Vinker, Lior Wolf et al.SIGGRAPH 2023 · 438 citations
- CLIPasso: semantically-aware object sketchingYael Vinker, Ehsan Pajouheshgar, Jessica Y. Bo, Roman Christian Bachmann et al.SIGGRAPH 2022 · 219 citations
- Nested Hierarchical Transformer: Towards Accurate, Data-Efficient and Interpretable Visual UnderstandingZizhao Zhang, Han Zhang, Long Zhao, Ting Chen et al.AAAI 2022 · 216 citations
- Iterative Prompt Learning for Unsupervised Backlit Image EnhancementZhexin Liang, Chongyi Li, Shangchen Zhou, Ruicheng Feng et al.ICCV 2023 · 196 citations
Builds on13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
Related papers
- VL-InterpreT: An Interactive Visualization Tool for Interpreting Vision-Language TransformersEstelle Aflalo, Meng Du, Shao-Yen Tseng, Yongfei Liu et al.CVPR 2022 · 34 citations
- LAVT: Language-Aware Vision Transformer for Referring Image SegmentationZhao Yang, Jiaqi Wang, Yansong Tang, Kai Chen et al.CVPR 2022 · 319 citations
- Transformer Interpretability Beyond Attention VisualizationHila Chefer, Shir Gur, Lior WolfCVPR 2021
- AttCAT: Explaining Transformers via Attentive Class Activation TokensYao Qiang, Deng Pan, Chengyin Li, Xin Li et al.NeurIPS 2022 · 66 citations
- VisQA: X-raying Vision and Language Reasoning in TransformersTheo Jaunet, Corentin Kervadec, Romain Vuillemot, Grigory Antipov et al.IEEE VIS 2021 · 30 citations
