A Mixture of h - 1 Heads is Better than h Heads
Hao Peng, Roy Schwartz, Dianqi Li, Noah A. Smith
摘要
Multi-head attentive neural architectures have achieved state-of-the-art results on a variety of natural language processing tasks. Evidence has shown that they are overparameterized; attention heads can be pruned without significant performance loss. In this work, we instead "reallocate" them-the model learns to activate different heads on different inputs. Drawing connections between multi-head attention and mixture of experts, we propose the mixture of attentive experts model (MAE). MAE is trained using a block coordinate descent algorithm that alternates between updating (1) the responsibilities of the experts and (2) their parameters. Experiments on machine translation and language modeling show that MAE outperforms strong baselines on both tasks. Particularly, on the WMT14 English to German translation dataset, MAE improves over "transformer-base" by 0.8 BLEU, with a comparable number of parameters. Our analysis shows that our model learns to specialize different experts to different inputs. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Random Feature AttentionHao Peng, Nikolaos Pappas, Dani Yogatama, Roy Schwartz 等ICLR 2021 · 被引用 425 次
- Conditionally Adaptive Multi-Task Learning: Improving Transfer Learning in NLP Using Fewer Parameters & Less DataJonathan Pilault, Amine Elhattami, Christopher J. PalICLR 2021 · 被引用 105 次
- Mixed SIGNals: Sign Language Production via a Mixture of Motion PrimitivesBen Saunders, Necati Cihan Camgöz, Richard BowdenICCV 2021 · 被引用 82 次
- BotMoE: Twitter Bot Detection with Community-Aware Mixtures of Modal-Specific ExpertsYuhan Liu, Zhaoxuan Tan, Heng Wang, Shangbin Feng 等SIGIR 2023 · 被引用 54 次
- SwitchHead: Accelerating Transformers with Mixture-of-Experts AttentionRóbert Csordás, Piotr Piekos, Kazuki Irie, Jürgen SchmidhuberNeurIPS 2024 · 被引用 46 次
它引用的顶会 Paper1
相关 Paper
- Mixture of Attention Heads: Selecting Attention Heads Per TokenXiaofeng Zhang, Yikang Shen, Zeyu Huang, Jie Zhou 等EMNLP 2022 · 被引用 23 次
- Cascaded Head-colliding AttentionLin Zheng, Zhiyong Wu, Lingpeng KongACL 2021
- Enlivening Redundant Heads in Multi-head Self-attention for Machine TranslationTianfu Zhang, Heyan Huang, Chong Feng, Longbing CaoEMNLP 2021 · 被引用 9 次
- Improving Transformers with Probabilistic Attention KeysTam Minh Nguyen, Tan Minh Nguyen, Dung D. Le, Duy Khuong Nguyen 等ICML 2022 · 被引用 38 次
- Pay Better Attention to Attention: Head Selection in Multilingual and Multi-Domain Sequence ModelingHongyu Gong, Yun Tang, Juan Miguel Pino, Xian LiNeurIPS 2021 · 被引用 15 次
