Attention Calibration for Transformer in Neural Machine Translation
Yu Lu, Jiali Zeng, Jiajun Zhang, Shuangzhi Wu, Mu Li
摘要
Attention mechanisms have achieved substantial improvements in neural machine translation by dynamically selecting relevant inputs for different predictions. However, recent studies have questioned the attention mechanisms' capability for discovering decisive inputs. In this paper, we propose to calibrate the attention weights by introducing a mask perturbation model that automatically evaluates each input's contribution to the model outputs. We increase the attention weights assigned to the indispensable tokens, whose removal leads to a dramatic performance decrease. The extensive experiments on the Transformer-based translation have demonstrated the effectiveness of our model. We further find that the calibrated attention weights are more uniform at lower layers to collect multiple information while more concentrated on the specific inputs at higher layers. Detailed analyses also show a great need for calibration in the attention weights with high entropy where the model is unconfident about its decision 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Attention-Aligned Transformer for Image CaptioningZhengcong FeiAAAI 2022 · 被引用 42 次
- SSPAttack: A Simple and Sweet Paradigm for Black-Box Hard-Label Textual Adversarial AttackHan Liu, Zhi Xu, Xiaotong Zhang, Xiaoming Xu 等AAAI 2023 · 被引用 31 次
- ZeroTuning: Unlocking the Initial Token's Power to Enhance Large Language Models Without TrainingFeijiang Han, Xiaodong Yu, Jianheng Tang, Delip Rao 等ICLR 2026 · 被引用 17 次
- Interpreting and Exploiting Functional Specialization in Multi-Head Attention under Multi-task LearningChong Li, Shaonan Wang, Yunhao Zhang, Jiajun Zhang 等EMNLP 2023 · 被引用 5 次
- SinkTrack: Attention Sink based Context Anchoring for Large Language ModelsXu Liu, Guikun Chen, Wenguan WangICLR 2026 · 被引用 4 次
它引用的顶会 Paper4
- Modeling Fluency and Faithfulness for Diverse Neural Machine TranslationYang Feng, Wanying Xie, Shuhao Gu, Chenze Shao 等AAAI 2020 · 被引用 28 次
- Towards Enhancing Faithfulness for Neural Machine TranslationRongxiang Weng, Heng Yu, Xiangpeng Wei, Weihua LuoEMNLP 2020 · 被引用 18 次
- Towards Transparent and Explainable Attention ModelsAkash Kumar Mohankumar, Preksha Nema, Sharan Narasimhan, Mitesh M. Khapra 等ACL 2020 · 被引用 11 次
- Analyzing the Source and Target Contributions to Predictions in Neural Machine TranslationElena Voita, Rico Sennrich, Ivan TitovACL 2021
相关 Paper
- Interpreting Positional Information in Perspective of Word OrderXilong Zhang, Ruochen Liu, Jin Liu, Xuefeng LiangACL 2023
- Attention is Not Only a Weight: Analyzing Transformers with Vector NormsGoro Kobayashi, Tatsuki Kuribayashi, Sho Yokoi, Kentaro InuiEMNLP 2020 · 被引用 138 次
- Recurrent Attention for Neural Machine TranslationJiali Zeng, Shuangzhi Wu, Yongjing Yin, Yufan Jiang 等EMNLP 2021
- Towards Opening the Black Box of Neural Machine Translation: Source and Target Interpretations of the TransformerJavier Ferrando, Gerard I. Gállego, Belen Alastruey, Carlos Escolano 等EMNLP 2022 · 被引用 18 次
- Why Attentions May Not Be Interpretable?Bing Bai, Jian Liang, Guanhua Zhang, Hao Li 等KDD 2021 · 被引用 51 次
