Contributions of Transformer Attention Heads in Multi- and Cross-lingual Tasks
Weicheng Ma, Kai Zhang, Renze Lou, Lili Wang, Soroush Vosoughi
Abstract
This paper studies the relative importance of attention heads in Transformer-based models to aid their interpretability in cross-lingual and multi-lingual tasks. Prior research has found that only a few attention heads are important in each mono-lingual Natural Language Processing (NLP) task and pruning the remaining heads leads to comparable or improved performance of the model. However, the impact of pruning attention heads is not yet clear in cross-lingual and multi-lingual tasks. Through extensive experiments, we show that (1) pruning a number of attention heads in a multilingual Transformer-based model has, in general, positive effects on its performance in cross-lingual and multi-lingual tasks and (2) the attention heads to be pruned can be ranked using gradients and identified with a few trial experiments. Our experiments focus on sequence labeling tasks, with potential applicability on other cross-lingual and multi-lingual tasks. For comprehensiveness, we examine two pre-trained multi-lingual models, namely multi-lingual BERT (mBERT) and XLM-R, on three tasks across 9 languages each. We also discuss the validity of our findings and their extensibility to truly resource-scarce languages and other task settings.
† Work done when interning at the Minds, Machines, and Society Lab at Dartmouth College. 1 We regard single-source machine translation as a monolingual task since the inputs to the models are mono-lingual.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e68beea6-7f84-478e-9953-1af337768492Cited by top-tier papers1
Ask how each one uses itBuilds on4
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- End-to-End Slot Alignment and Recognition for Cross-Lingual NLUWeijia Xu, Batool Haider, Saab MansourEMNLP 2020 · 109 citations
- Single-/Multi-Source Cross-Lingual NER via Teacher-Student Learning on Unlabeled Data in Target LanguageQianhui Wu, Zijia Lin, Börje Karlsson, Jianguang Lou et al.ACL 2020 · 59 citations
- Unsupervised Cross-Lingual Part-of-Speech Tagging for Truly Low-Resource ScenariosRamy Eskander, Smaranda Muresan, Michael CollinsEMNLP 2020 · 16 citations
Related papers
- Focusing on Language: Revealing and Exploiting Language Attention Heads in Multilingual Large Language ModelsXin Liu, Qiyang Song, Qihang Zhou, Haichao Du et al.AAAI 2026
- Roles and Utilization of Attention Heads in Transformer-based Neural Language ModelsJae-young Jo, Sung-Hyon MyaengACL 2020 · 32 citations
- Losing Heads in the Lottery: Pruning Transformer Attention in Neural Machine TranslationMaximiliana Behnke, Kenneth HeafieldEMNLP 2020 · 50 citations
- The Stem Cell Hypothesis: Dilemma behind Multi-Task Learning with Transformer EncodersHan He, Jinho D. ChoiEMNLP 2021 · 111 citations
- Multilingual Pre-training with Universal Dependency LearningKailai Sun, Zuchao Li, Hai ZhaoNeurIPS 2021 · 11 citations
