SEAT: Stable and Explainable Attention
Lijie Hu, Yixin Liu, Ninghao Liu, Mengdi Huai, Lichao Sun, Di Wang
Abstract
Currently, attention mechanism becomes a standard fixture in most state-of-the-art natural language processing (NLP) models, not only due to outstanding performance it could gain, but also due to plausible innate explanation for the behaviors of neural architectures it provides, which is notoriously difficult to analyze. However, recent studies show that attention is unstable against randomness and perturbations during training or testing, such as random seeds and slight perturbation of embedding vectors, which impedes it from becoming a faithful explanation tool. Thus, a natural question is whether we can find some substitute of the current attention which is more stable and could keep the most important characteristics on explanation and prediction of attention. In this paper, to resolve the problem, we provide a first rigorous definition of such alternate namely SEAT (Stable and Explainable Attention). Specifically, a SEAT should has the following three properties: (1) Its prediction distribution is enforced to be close to the distribution based on the vanilla attention; (2) Its top-k indices have large overlaps with those of the vanilla attention; (3) It is robust w.r.t perturbations, i.e., any slight perturbation on SEAT will not change the prediction distribution too much, which implicitly indicates that it is stable to randomness and perturbations. Moreover we propose a method to get a SEAT, which could be considered as an ad hoc modification for the canonical attention. Finally, through intensive experiments on various datasets, we compare our SEAT with other baseline methods using RNN, BiLSTM and BERT architectures via six different evaluation metrics for model interpretation, stability and accuracy. Results show that SEAT is more stable against different perturbations and randomness while also keeps the explainability of attention, which indicates it is a more faithful explanation. Moreover, compared with vanilla attention, there is almost no utility (accuracy) degradation for SEAT.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- Faithful Vision-Language Interpretation via Concept Bottleneck ModelsSongning Lai, Lijie Hu, Junxiao Wang, Laure Berti-Équille et al.ICLR 2024 · 42 citations
- Towards Understanding and Enhancing Robustness of Deep Learning Models against Malicious Unlearning AttacksWei Qian, Chenxu Zhao, Wei Le, Meiyi Ma et al.KDD 2023 · 38 citations
- SATO: Stable Text-to-Motion FrameworkWenshuo Chen, Hongru Xiao, Erhang Zhang, Lijie Hu et al.ACM MM 2024 · 17 citations
- Improving Interpretation Faithfulness for Vision TransformersLijie Hu, Yixin Liu, Ninghao Liu, Mengdi Huai et al.ICML 2024 · 12 citations
- Towards Multi-dimensional Explanation Alignment for Medical ClassificationLijie Hu, Songning Lai, Wenshuo Chen, Hongru Xiao et al.NeurIPS 2024 · 8 citations
Builds on3
- ERASER: A Benchmark to Evaluate Rationalized NLP ModelsJay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric P. Lehman et al.ACL 2020 · 36 citations
- On the Sensitivity and Stability of Model Interpretations in NLPFan Yin, Zhouxing Shi, Cho-Jui Hsieh, Kai-Wei ChangACL 2022 · 35 citations
- Towards Transparent and Explainable Attention ModelsAkash Kumar Mohankumar, Preksha Nema, Sharan Narasimhan, Mitesh M. Khapra et al.ACL 2020 · 11 citations
Related papers
- Is Attention Explanation? An Introduction to the DebateAdrien Bibal, Rémi Cardon, David Alfter, Rodrigo Wilkens et al.ACL 2022
- Generating Hierarchical Explanations on Text Classification via Feature Interaction DetectionHanjie Chen, Guangtao Zheng, Yangfeng JiACL 2020 · 85 citations
- Improving the Faithfulness of Attention-based Explanations with Task-specific Information for Text ClassificationGeorge Chrysostomou, Nikolaos AletrasACL 2021
- XAI for Transformers: Better Explanations through Conservative PropagationAmeen Ali, Thomas Schnake, Oliver Eberle, Grégoire Montavon et al.ICML 2022 · 144 citations
- Attention-based Interpretability with Concept TransformersMattia Rigotti, Christoph Miksovic, Ioana Giurgiu, Thomas Gschwind et al.ICLR 2022 · 77 citations
