Cascaded Head-colliding Attention
Lin Zheng, Zhiyong Wu, Lingpeng Kong
摘要
Transformers have advanced the field of natural language processing (NLP) in many ways. At the heart of the Transformer architecture is the multi-head attention (MHA) mechanism which models pairwise interactions between the elements of the sequence. Despite its massive success, the current framework ignores interactions among different heads, leading to the problem that many of the heads are redundant in practice, which underutilizes the capacity of the model. To improve parameter efficiency, we re-formulate the MHA as a latent variable model from a probabilistic perspective. We present cascaded head-colliding attention (CODA) which explicitly models the interactions between attention heads through a hierarchical variational distribution. We conduct extensive experiments and demonstrate that CODA outperforms the transformer baseline, by 0.6 perplexity on Wikitext-103 in language modeling, and by 0.6 BLEU on WMT14 EN-DE in machine translation, due to its improvements on the parameter efficiency. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Perturbed Masking: Parameter-free Probing for Analyzing and Interpreting BERTZhiyong Wu, Yun Chen, Ben Kao, Qun LiuACL 2020 · 被引用 158 次
- Bayesian Attention ModulesXinjie Fan, Shujian Zhang, Bo Chen, Mingyuan ZhouNeurIPS 2020 · 被引用 78 次
- Hard-Coded Gaussian Attention for Neural Machine TranslationWeiqiu You, Simeng Sun, Mohit IyyerACL 2020 · 被引用 55 次
- Losing Heads in the Lottery: Pruning Transformer Attention in Neural Machine TranslationMaximiliana Behnke, Kenneth HeafieldEMNLP 2020 · 被引用 50 次
相关 Paper
- A Mixture of h - 1 Heads is Better than h HeadsHao Peng, Roy Schwartz, Dianqi Li, Noah A. SmithACL 2020 · 被引用 25 次
- Improving Transformers with Dynamically Composable Multi-Head AttentionDa Xiao, Qingye Meng, Shengping Li, Xingyuan YuanICML 2024 · 被引用 8 次
- Improving Transformers with Probabilistic Attention KeysTam Minh Nguyen, Tan Minh Nguyen, Dung D. Le, Duy Khuong Nguyen 等ICML 2022 · 被引用 38 次
- Multi-Unit Transformers for Neural Machine TranslationJianhao Yan, Fandong Meng, Jie ZhouEMNLP 2020 · 被引用 21 次
- Recurrent Attention for Neural Machine TranslationJiali Zeng, Shuangzhi Wu, Yongjing Yin, Yufan Jiang 等EMNLP 2021
