Hard-Coded Gaussian Attention for Neural Machine Translation
Weiqiu You, Simeng Sun, Mohit Iyyer
摘要
Recent work has questioned the importance of the Transformer's multi-headed attention for achieving high translation quality. We push further in this direction by developing a "hardcoded" attention variant without any learned parameters. Surprisingly, replacing all learned self-attention heads in the encoder and decoder with fixed, input-agnostic Gaussian distributions minimally impacts BLEU scores across four different language pairs. However, additionally hard-coding cross attention (which connects the decoder to the encoder) significantly lowers BLEU, suggesting that it is more important than self-attention. Much of this BLEU drop can be recovered by adding just a single learned cross attention head to an otherwise hard-coded Transformer. Taken as a whole, our results offer insight into which components of the Transformer are actually important, which we hope will guide future work into the development of simpler and more efficient attention-based models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Random Feature AttentionHao Peng, Nikolaos Pappas, Dani Yogatama, Roy Schwartz 等ICLR 2021 · 被引用 425 次
- Cross-Attention is All You Need: Adapting Pretrained Transformers for Machine TranslationMozhdeh Gheini, Xiang Ren, Jonathan MayEMNLP 2021 · 被引用 133 次
- A Trainable Optimal Transport Embedding for Feature Aggregation and its Relationship to AttentionGrégoire Mialon, Dexiong Chen, Alexandre d'Aspremont, Julien MairalICLR 2021 · 被引用 71 次
- Sparse and Continuous Attention MechanismsAndré F. T. Martins, António Farinhas, Marcos V. Treviso, Vlad Niculae 等NeurIPS 2020 · 被引用 55 次
- Adversarial Self-Attention for Language UnderstandingHongqiu Wu, Ruixue Ding, Hai Zhao, Pengjun Xie 等AAAI 2023 · 被引用 19 次
相关 Paper
- Losing Heads in the Lottery: Pruning Transformer Attention in Neural Machine TranslationMaximiliana Behnke, Kenneth HeafieldEMNLP 2020 · 被引用 50 次
- Recurrent Attention for Neural Machine TranslationJiali Zeng, Shuangzhi Wu, Yongjing Yin, Yufan Jiang 等EMNLP 2021
- Scaling Laws for Neural Machine TranslationBehrooz Ghorbani, Orhan Firat, Markus Freitag, Ankur Bapna 等ICLR 2022 · 被引用 130 次
- An Efficient Transformer Decoder with Compressed Sub-layersYanyang Li, Ye Lin, Tong Xiao, Jingbo ZhuAAAI 2021 · 被引用 32 次
- Improving Transformers with Probabilistic Attention KeysTam Minh Nguyen, Tan Minh Nguyen, Dung D. Le, Duy Khuong Nguyen 等ICML 2022 · 被引用 38 次
