Attention-based Interpretability with Concept Transformers
Mattia Rigotti, Christoph Miksovic, Ioana Giurgiu, Thomas Gschwind, Paolo Scotton
Abstract
Attention is a mechanism that has been instrumental in driving remarkable performance gains of deep neural network models in a host of visual, NLP and multimodal tasks.One additional notable aspect of attention is that it conveniently exposes the ``reasoning'' behind each particular output generated by the model.Specifically, attention scores over input regions or intermediate features have been interpreted as a measure of the contribution of the attended element to the model inference.While the debate in regard to the interpretability of attention is still not settled, researchers have pointed out the existence of architectures and scenarios that afford a meaningful interpretation of the attention mechanism.Here we propose the generalization of attention from low-level input features to high-level concepts as a mechanism to ensure the interpretability of attention scores within a given application domain.In particular, we design the ConceptTransformer, a deep learning module that exposes explanations of the output of a model in which it is embedded in terms of attention over user-defined high-level concepts.Such explanations are plausible (i.e. convincing to the human user) and faithful (i.e. truly reflective of the reasoning process of the model).Plausibility of such explanations is obtained by construction by training the attention heads to conform with known relations between inputs, concepts and outputs dictated by domain knowledge.Faithfulness is achieved by design by enforcing a linear relation between the transformer value vectors that represent the concepts and their contribution to the classification log-probabilities.We validate our ConceptTransformer module on established explainability benchmarks and show how it can be used to infuse domain knowledge into classifiers to improve accuracy, and conversely to extract concept-based explanations of classification outputs. Code to reproduce our results is available at: https://github.com/ibm/concept_transformer.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 67ec9e9d-863f-4833-ba70-aa88a62512f7Cited by top-tier papers15
- MICA: Towards Explainable Skin Lesion Diagnosis via Multi-Level Image-Concept AlignmentYequan Bie, Luyang Luo, Hao ChenAAAI 2024 · 28 citations
- A Simple Interpretable Transformer for Fine-Grained Image Classification and AnalysisDipanjyoti Paul, Arpita Chowdhury, Xinqi Xiong, Feng-Ju Chang et al.ICLR 2024 · 27 citations
- Towards Modeling Uncertainties of Self-Explaining Neural Networks via Conformal PredictionWei Qian, Chenxu Zhao, Yangyi Li, Fenglong Ma et al.AAAI 2024 · 14 citations
- Towards Compositionality in Concept LearningAdam Stein, Aaditya Naik, Yinjun Wu, Mayur Naik et al.ICML 2024 · 11 citations
- Mitigating the Effect of Incidental Correlations on Part-based LearningGaurav Bhatt, Deepayan Das, Leonid Sigal, Vineeth N. BalasubramanianNeurIPS 2023 · 7 citations
Related papers
- Generic Attention-model Explainability for Interpreting Bi-Modal and Encoder-Decoder TransformersHila Chefer, Shir Gur, Lior WolfICCV 2021 · 451 citations
- FaCT: Faithful Concept Traces for Explaining Neural Network DecisionsAmin Parchami-Araghi, Sukrut Rao, Jonas Fischer, Bernt SchieleNeurIPS 2025 · 1 citation
- Towards Transparent and Explainable Attention ModelsAkash Kumar Mohankumar, Preksha Nema, Sharan Narasimhan, Mitesh M. Khapra et al.ACL 2020 · 11 citations
- Is Attention Explanation? An Introduction to the DebateAdrien Bibal, Rémi Cardon, David Alfter, Rodrigo Wilkens et al.ACL 2022
- AttCAT: Explaining Transformers via Attentive Class Activation TokensYao Qiang, Deng Pan, Chengyin Li, Xin Li et al.NeurIPS 2022 · 66 citations
