Finding the Pillars of Strength for Multi-Head Attention
Jinjie Ni, Rui Mao, Zonglin Yang, Han Lei, Erik Cambria
Abstract
Recent studies have revealed some issues of Multi-Head Attention (MHA), e.g., redundancy and over-parameterization. Specifically, the heads of MHA were originally designed to attend to information from different representation subspaces, whereas prior studies found that some attention heads likely learn similar features and can be pruned without harming performance. Inspired by the minimum-redundancy feature selection, we assume that focusing on the most representative and distinctive features with minimum resources can mitigate the above issues and lead to more effective and efficient MHAs. In particular, we propose Grouped Head Attention, trained with a self-supervised group constraint that group attention heads, where each group focuses on an essential but distinctive feature subset. We additionally propose a Voting-to-Stay procedure to remove redundant heads, thus achieving a transformer with lighter weights. Extensive experiments are consistent with our hypothesis. Moreover, our method achieves significant performance gains on three well-established tasks while considerably compressing parameters.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext db7de6d1-caa8-4e2e-aefb-2d8efd9b8742Cited by top-tier papers1
Ask how each one uses itBuilds on2
Related papers
- Curse of High Dimensionality Issue in Transformer for Long Context ModelingShuhai Zhang, Zeng You, Yaofo Chen, Zhiquan Wen et al.ICML 2025
- DHA: Learning Decoupled-Head Attention from Transformer Checkpoints via Adaptive Heads FusionYilong Chen, Linhao Zhang, Junyuan Shang, Zhenyu Zhang et al.NeurIPS 2024 · 12 citations
- Enlivening Redundant Heads in Multi-head Self-attention for Machine TranslationTianfu Zhang, Heyan Huang, Chong Feng, Longbing CaoEMNLP 2021 · 9 citations
- SAS: Simulated Attention ScoreChuanyang Zheng, Jiankai Sun, Yihang Gao, Yuehao Wang et al.NeurIPS 2025 · 3 citations
- Improving Transformers with Dynamically Composable Multi-Head AttentionDa Xiao, Qingye Meng, Shengping Li, Xingyuan YuanICML 2024 · 8 citations
