Enlivening Redundant Heads in Multi-head Self-attention for Machine Translation
Tianfu Zhang, Heyan Huang, Chong Feng, Longbing Cao
Abstract
Multi-head self-attention recently attracts enormous interest owing to its specialized functions, significant parallelizable computation, and flexible extensibility. However, very recent empirical studies show that some self-attention heads make little contribution and can be pruned as redundant heads. This work takes a novel perspective of identifying and then vitalizing redundant heads. We propose a redundant head enlivening (RHE) method to precisely identify redundant heads, and then vitalize their potential by learning syntactic relations and prior knowledge in the text without sacrificing the roles of important heads. Two novel syntax-enhanced attention (SEA) mechanisms: a dependency mask bias and a relative local-phrasal position bias, are introduced to revise self-attention distributions for syntactic enhancement in machine translation. The importance of individual heads is dynamically evaluated during the redundant heads identification, on which we apply SEA to vitalize redundant heads while maintaining the strength of important heads. Experimental results on widely adopted WMT14 and WMT16 English to German and English to Czech language machine translation validate the RHE effectiveness.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 64de45ea-8c53-4386-82df-5fbee18e485eCited by top-tier papers3
- Gloss Semantic-Enhanced Network with Online Back-Translation for Sign Language ProductionShengeng Tang, Richang Hong, Dan Guo, Meng WangACM MM 2022 · 44 citations
- Budgeted Training for Vision TransformerZhuofan Xia, Xuran Pan, Xuan Jin, Yuan He et al.ICLR 2023
- Dynamic Diffusion TransformerWangbo Zhao, Yizeng Han, Jiasheng Tang, Kai Wang et al.ICLR 2025
Related papers
- A Mixture of h - 1 Heads is Better than h HeadsHao Peng, Roy Schwartz, Dianqi Li, Noah A. SmithACL 2020 · 25 citations
- Finding the Pillars of Strength for Multi-Head AttentionJinjie Ni, Rui Mao, Zonglin Yang, Han Lei et al.ACL 2023 · 9 citations
- Contributions of Transformer Attention Heads in Multi- and Cross-lingual TasksWeicheng Ma, Kai Zhang, Renze Lou, Lili Wang et al.ACL 2021
- EIT: Enhanced Interactive TransformerTong Zheng, Bei Li, Huiwen Bao, Tong Xiao et al.ACL 2024
- Losing Heads in the Lottery: Pruning Transformer Attention in Neural Machine TranslationMaximiliana Behnke, Kenneth HeafieldEMNLP 2020 · 50 citations
