Attention-based Interpretability with Concept Transformers
Mattia Rigotti, Christoph Miksovic, Ioana Giurgiu, Thomas Gschwind, Paolo Scotton
摘要
Attention is a mechanism that has been instrumental in driving remarkable performance gains of deep neural network models in a host of visual, NLP and multimodal tasks.One additional notable aspect of attention is that it conveniently exposes the ``reasoning'' behind each particular output generated by the model.Specifically, attention scores over input regions or intermediate features have been interpreted as a measure of the contribution of the attended element to the model inference.While the debate in regard to the interpretability of attention is still not settled, researchers have pointed out the existence of architectures and scenarios that afford a meaningful interpretation of the attention mechanism.Here we propose the generalization of attention from low-level input features to high-level concepts as a mechanism to ensure the interpretability of attention scores within a given application domain.In particular, we design the ConceptTransformer, a deep learning module that exposes explanations of the output of a model in which it is embedded in terms of attention over user-defined high-level concepts.Such explanations are plausible (i.e. convincing to the human user) and faithful (i.e. truly reflective of the reasoning process of the model).Plausibility of such explanations is obtained by construction by training the attention heads to conform with known relations between inputs, concepts and outputs dictated by domain knowledge.Faithfulness is achieved by design by enforcing a linear relation between the transformer value vectors that represent the concepts and their contribution to the classification log-probabilities.We validate our ConceptTransformer module on established explainability benchmarks and show how it can be used to infuse domain knowledge into classifiers to improve accuracy, and conversely to extract concept-based explanations of classification outputs. Code to reproduce our results is available at: https://github.com/ibm/concept_transformer.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper15
- MICA: Towards Explainable Skin Lesion Diagnosis via Multi-Level Image-Concept AlignmentYequan Bie, Luyang Luo, Hao ChenAAAI 2024 · 被引用 28 次
- A Simple Interpretable Transformer for Fine-Grained Image Classification and AnalysisDipanjyoti Paul, Arpita Chowdhury, Xinqi Xiong, Feng-Ju Chang 等ICLR 2024 · 被引用 27 次
- Towards Modeling Uncertainties of Self-Explaining Neural Networks via Conformal PredictionWei Qian, Chenxu Zhao, Yangyi Li, Fenglong Ma 等AAAI 2024 · 被引用 14 次
- Towards Compositionality in Concept LearningAdam Stein, Aaditya Naik, Yinjun Wu, Mayur Naik 等ICML 2024 · 被引用 11 次
- Mitigating the Effect of Incidental Correlations on Part-based LearningGaurav Bhatt, Deepayan Das, Leonid Sigal, Vineeth N. BalasubramanianNeurIPS 2023 · 被引用 7 次
相关 Paper
- Generic Attention-model Explainability for Interpreting Bi-Modal and Encoder-Decoder TransformersHila Chefer, Shir Gur, Lior WolfICCV 2021 · 被引用 451 次
- FaCT: Faithful Concept Traces for Explaining Neural Network DecisionsAmin Parchami-Araghi, Sukrut Rao, Jonas Fischer, Bernt SchieleNeurIPS 2025 · 被引用 1 次
- Towards Transparent and Explainable Attention ModelsAkash Kumar Mohankumar, Preksha Nema, Sharan Narasimhan, Mitesh M. Khapra 等ACL 2020 · 被引用 11 次
- Is Attention Explanation? An Introduction to the DebateAdrien Bibal, Rémi Cardon, David Alfter, Rodrigo Wilkens 等ACL 2022
- AttCAT: Explaining Transformers via Attentive Class Activation TokensYao Qiang, Deng Pan, Chengyin Li, Xin Li 等NeurIPS 2022 · 被引用 66 次
