Cascaded Head-colliding Attention
Lin Zheng, Zhiyong Wu, Lingpeng Kong
Abstract
Transformers have advanced the field of natural language processing (NLP) in many ways. At the heart of the Transformer architecture is the multi-head attention (MHA) mechanism which models pairwise interactions between the elements of the sequence. Despite its massive success, the current framework ignores interactions among different heads, leading to the problem that many of the heads are redundant in practice, which underutilizes the capacity of the model. To improve parameter efficiency, we re-formulate the MHA as a latent variable model from a probabilistic perspective. We present cascaded head-colliding attention (CODA) which explicitly models the interactions between attention heads through a hierarchical variational distribution. We conduct extensive experiments and demonstrate that CODA outperforms the transformer baseline, by 0.6 perplexity on Wikitext-103 in language modeling, and by 0.6 BLEU on WMT14 EN-DE in machine translation, due to its improvements on the parameter efficiency. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 74b055b2-3a2a-4b7b-b995-11e786a22cc6Builds on7
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Perturbed Masking: Parameter-free Probing for Analyzing and Interpreting BERTZhiyong Wu, Yun Chen, Ben Kao, Qun LiuACL 2020 · 158 citations
- Bayesian Attention ModulesXinjie Fan, Shujian Zhang, Bo Chen, Mingyuan ZhouNeurIPS 2020 · 78 citations
- Hard-Coded Gaussian Attention for Neural Machine TranslationWeiqiu You, Simeng Sun, Mohit IyyerACL 2020 · 55 citations
- Losing Heads in the Lottery: Pruning Transformer Attention in Neural Machine TranslationMaximiliana Behnke, Kenneth HeafieldEMNLP 2020 · 50 citations
Related papers
- A Mixture of h - 1 Heads is Better than h HeadsHao Peng, Roy Schwartz, Dianqi Li, Noah A. SmithACL 2020 · 25 citations
- Improving Transformers with Dynamically Composable Multi-Head AttentionDa Xiao, Qingye Meng, Shengping Li, Xingyuan YuanICML 2024 · 8 citations
- Improving Transformers with Probabilistic Attention KeysTam Minh Nguyen, Tan Minh Nguyen, Dung D. Le, Duy Khuong Nguyen et al.ICML 2022 · 38 citations
- Multi-Unit Transformers for Neural Machine TranslationJianhao Yan, Fandong Meng, Jie ZhouEMNLP 2020 · 21 citations
- Recurrent Attention for Neural Machine TranslationJiali Zeng, Shuangzhi Wu, Yongjing Yin, Yufan Jiang et al.EMNLP 2021
