Learning to Deceive with Attention-Based Explanations
Danish Pruthi, Mansi Gupta, Bhuwan Dhingra, Graham Neubig, Zachary C. Lipton
摘要
Attention mechanisms are ubiquitous components in neural architectures applied to natural language processing. In addition to yielding gains in predictive accuracy, attention weights are often claimed to confer interpretability, purportedly useful both for providing insights to practitioners and for explaining why a model makes its decisions to stakeholders. We call the latter use of attention mechanisms into question by demonstrating a simple method for training models to produce deceptive attention masks. Our method diminishes the total weight assigned to designated impermissible tokens, even when the models can be shown to nevertheless rely on these features to drive predictions. Across multiple models and tasks, our approach manipulates attention weights while paying surprisingly little cost in accuracy. Through a human study, we show that our manipulated attention-based explanations deceive people into thinking that predictions from a model biased against gender minorities do not rely on the gender. Consequently, our results cast doubt on attention's reliability as a tool for auditing algorithms in the context of fairness and accountability. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper42
- BERTology Meets Biology: Interpreting Attention in Protein Language ModelsJesse Vig, Ali Madani, Lav R. Varshney, Caiming Xiong 等ICLR 2021 · 被引用 357 次
- On Identifiability in TransformersGino Brunner, Yang Liu, Damian Pascual, Oliver Richter 等ICLR 2020 · 被引用 210 次
- Do Feature Attribution Methods Correctly Attribute Features?Yilun Zhou, Serena Booth, Marco Túlio Ribeiro, Julie ShahAAAI 2022 · 被引用 167 次
- Attention is Not Only a Weight: Analyzing Transformers with Vector NormsGoro Kobayashi, Tatsuki Kuribayashi, Sho Yokoi, Kentaro InuiEMNLP 2020 · 被引用 138 次
- Explanations, Fairness, and Appropriate Reliance in Human-AI Decision-MakingJakob Schoeffer, Maria De-Arteaga, Niklas KühlCHI 2024 · 被引用 67 次
相关 Paper
- Improving the Faithfulness of Attention-based Explanations with Task-specific Information for Text ClassificationGeorge Chrysostomou, Nikolaos AletrasACL 2021
- Fooling SHAP with Stealthily Biased SamplingGabriel Laberge, Ulrich Aïvodji, Satoshi Hara, Mario Marchand 等ICLR 2023 · 被引用 3 次
- Debiasing Pretrained Text Encoders by Paying Attention to Paying AttentionYacine Gaci, Boualem Benatallah, Fabio Casati, Khalid BenabdeslemEMNLP 2022 · 被引用 12 次
- Towards Transparent and Explainable Attention ModelsAkash Kumar Mohankumar, Preksha Nema, Sharan Narasimhan, Mitesh M. Khapra 等ACL 2020 · 被引用 11 次
- Why Attentions May Not Be Interpretable?Bing Bai, Jian Liang, Guanhua Zhang, Hao Li 等KDD 2021 · 被引用 51 次
