Roles and Utilization of Attention Heads in Transformer-based Neural Language Models
Jae-young Jo, Sung-Hyon Myaeng
Abstract
Sentence encoders based on the transformer architecture have shown promising results on various natural language tasks. The main impetus lies in the pre-trained neural language models that capture long-range dependencies among words, owing to multi-head attention that is unique in the architecture. However, little is known for how linguistic properties are processed, represented, and utilized for downstream tasks among hundreds of attention heads inside the pre-trained transformerbased model. For the initial goal of examining the roles of attention heads in handling a set of linguistic features, we conducted a set of experiments with ten probing tasks and three downstream tasks on four pre-trained transformer families (GPT, GPT2, BERT, and ELECTRA). Meaningful insights are shown through the lens of heat map visualization and utilized to propose a relatively simple sentence representation method that takes advantage of most influential attention heads, resulting in additional performance improvements on the downstream tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers12
- ATMAN: Understanding Transformer Predictions Through Memory Efficient Attention ManipulationBjörn Deiseroth, Mayukh Deb, Samuel Weinbach, Manuel Brack et al.NeurIPS 2023 · 45 citations
- FIND: Human-in-the-Loop Debugging Deep Text ClassifiersPiyawat Lertvittayakumjorn, Lucia Specia, Francesca ToniEMNLP 2020 · 33 citations
- Artificial Text Detection via Examining the Topology of Attention MapsLaida Kushnareva, Daniil Cherniavskii, Vladislav Mikhailov, Ekaterina Artemova et al.EMNLP 2021 · 27 citations
- GPT-D: Inducing Dementia-related Linguistic Anomalies by Deliberate Degradation of Artificial Neural Language ModelsChangye Li, David S. Knopman, Weizhe Xu, Trevor Cohen et al.ACL 2022 · 24 citations
- Ultra-High Dimensional Sparse Representations with Binarization for Efficient Text RetrievalKyoungrok Jang, Junmo Kang, Giwon Hong, Sung-Hyon Myaeng et al.EMNLP 2021 · 13 citations
Builds on1
Related papers
- Less Mature is More Adaptable for Sentence-level Language ModelingAbhilasha Sancheti, David Dale, Artyom Kozhevnikov, Maha ElbayadACL 2025
- Contributions of Transformer Attention Heads in Multi- and Cross-lingual TasksWeicheng Ma, Kai Zhang, Renze Lou, Lili Wang et al.ACL 2021
- Analyzing Individual Neurons in Pre-trained Language ModelsNadir Durrani, Hassan Sajjad, Fahim Dalvi, Yonatan BelinkovEMNLP 2020 · 5 citations
- Map of Encoders - Mapping Sentence Encoders using Quantum Relative EntropyGaifan Zhang, Danushka BollegalaACL 2026
- How does BERT's attention change when you fine-tune? An analysis methodology and a case study in negation scopeYiyun Zhao, Steven BethardACL 2020 · 35 citations
