Attention Calibration for Transformer in Neural Machine Translation
Yu Lu, Jiali Zeng, Jiajun Zhang, Shuangzhi Wu, Mu Li
Abstract
Attention mechanisms have achieved substantial improvements in neural machine translation by dynamically selecting relevant inputs for different predictions. However, recent studies have questioned the attention mechanisms' capability for discovering decisive inputs. In this paper, we propose to calibrate the attention weights by introducing a mask perturbation model that automatically evaluates each input's contribution to the model outputs. We increase the attention weights assigned to the indispensable tokens, whose removal leads to a dramatic performance decrease. The extensive experiments on the Transformer-based translation have demonstrated the effectiveness of our model. We further find that the calibrated attention weights are more uniform at lower layers to collect multiple information while more concentrated on the specific inputs at higher layers. Detailed analyses also show a great need for calibration in the attention weights with high entropy where the model is unconfident about its decision 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8bef94a4-d801-4223-b80e-93c460cc5bf0Cited by top-tier papers7
- Attention-Aligned Transformer for Image CaptioningZhengcong FeiAAAI 2022 · 42 citations
- SSPAttack: A Simple and Sweet Paradigm for Black-Box Hard-Label Textual Adversarial AttackHan Liu, Zhi Xu, Xiaotong Zhang, Xiaoming Xu et al.AAAI 2023 · 31 citations
- ZeroTuning: Unlocking the Initial Token's Power to Enhance Large Language Models Without TrainingFeijiang Han, Xiaodong Yu, Jianheng Tang, Delip Rao et al.ICLR 2026 · 17 citations
- Interpreting and Exploiting Functional Specialization in Multi-Head Attention under Multi-task LearningChong Li, Shaonan Wang, Yunhao Zhang, Jiajun Zhang et al.EMNLP 2023 · 5 citations
- SinkTrack: Attention Sink based Context Anchoring for Large Language ModelsXu Liu, Guikun Chen, Wenguan WangICLR 2026 · 4 citations
Builds on4
- Modeling Fluency and Faithfulness for Diverse Neural Machine TranslationYang Feng, Wanying Xie, Shuhao Gu, Chenze Shao et al.AAAI 2020 · 28 citations
- Towards Enhancing Faithfulness for Neural Machine TranslationRongxiang Weng, Heng Yu, Xiangpeng Wei, Weihua LuoEMNLP 2020 · 18 citations
- Towards Transparent and Explainable Attention ModelsAkash Kumar Mohankumar, Preksha Nema, Sharan Narasimhan, Mitesh M. Khapra et al.ACL 2020 · 11 citations
- Analyzing the Source and Target Contributions to Predictions in Neural Machine TranslationElena Voita, Rico Sennrich, Ivan TitovACL 2021
Related papers
- Interpreting Positional Information in Perspective of Word OrderXilong Zhang, Ruochen Liu, Jin Liu, Xuefeng LiangACL 2023
- Attention is Not Only a Weight: Analyzing Transformers with Vector NormsGoro Kobayashi, Tatsuki Kuribayashi, Sho Yokoi, Kentaro InuiEMNLP 2020 · 138 citations
- Recurrent Attention for Neural Machine TranslationJiali Zeng, Shuangzhi Wu, Yongjing Yin, Yufan Jiang et al.EMNLP 2021
- Towards Opening the Black Box of Neural Machine Translation: Source and Target Interpretations of the TransformerJavier Ferrando, Gerard I. Gállego, Belen Alastruey, Carlos Escolano et al.EMNLP 2022 · 18 citations
- Why Attentions May Not Be Interpretable?Bing Bai, Jian Liang, Guanhua Zhang, Hao Li et al.KDD 2021 · 51 citations
