A Mixture of h - 1 Heads is Better than h Heads
Hao Peng, Roy Schwartz, Dianqi Li, Noah A. Smith
Abstract
Multi-head attentive neural architectures have achieved state-of-the-art results on a variety of natural language processing tasks. Evidence has shown that they are overparameterized; attention heads can be pruned without significant performance loss. In this work, we instead "reallocate" them-the model learns to activate different heads on different inputs. Drawing connections between multi-head attention and mixture of experts, we propose the mixture of attentive experts model (MAE). MAE is trained using a block coordinate descent algorithm that alternates between updating (1) the responsibilities of the experts and (2) their parameters. Experiments on machine translation and language modeling show that MAE outperforms strong baselines on both tasks. Particularly, on the WMT14 English to German translation dataset, MAE improves over "transformer-base" by 0.8 BLEU, with a comparable number of parameters. Our analysis shows that our model learns to specialize different experts to different inputs. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 80c1fbb4-1294-45e0-a100-fdb866d4e4edCited by top-tier papers13
- Random Feature AttentionHao Peng, Nikolaos Pappas, Dani Yogatama, Roy Schwartz et al.ICLR 2021 · 425 citations
- Conditionally Adaptive Multi-Task Learning: Improving Transfer Learning in NLP Using Fewer Parameters & Less DataJonathan Pilault, Amine Elhattami, Christopher J. PalICLR 2021 · 105 citations
- Mixed SIGNals: Sign Language Production via a Mixture of Motion PrimitivesBen Saunders, Necati Cihan Camgöz, Richard BowdenICCV 2021 · 82 citations
- BotMoE: Twitter Bot Detection with Community-Aware Mixtures of Modal-Specific ExpertsYuhan Liu, Zhaoxuan Tan, Heng Wang, Shangbin Feng et al.SIGIR 2023 · 54 citations
- SwitchHead: Accelerating Transformers with Mixture-of-Experts AttentionRóbert Csordás, Piotr Piekos, Kazuki Irie, Jürgen SchmidhuberNeurIPS 2024 · 46 citations
Builds on1
Related papers
- Mixture of Attention Heads: Selecting Attention Heads Per TokenXiaofeng Zhang, Yikang Shen, Zeyu Huang, Jie Zhou et al.EMNLP 2022 · 23 citations
- Cascaded Head-colliding AttentionLin Zheng, Zhiyong Wu, Lingpeng KongACL 2021
- Enlivening Redundant Heads in Multi-head Self-attention for Machine TranslationTianfu Zhang, Heyan Huang, Chong Feng, Longbing CaoEMNLP 2021 · 9 citations
- Improving Transformers with Probabilistic Attention KeysTam Minh Nguyen, Tan Minh Nguyen, Dung D. Le, Duy Khuong Nguyen et al.ICML 2022 · 38 citations
- Pay Better Attention to Attention: Head Selection in Multilingual and Multi-Domain Sequence ModelingHongyu Gong, Yun Tang, Juan Miguel Pino, Xian LiNeurIPS 2021 · 15 citations
