Roles and Utilization of Attention Heads in Transformer-based Neural Language Models
Jae-young Jo, Sung-Hyon Myaeng
摘要
Sentence encoders based on the transformer architecture have shown promising results on various natural language tasks. The main impetus lies in the pre-trained neural language models that capture long-range dependencies among words, owing to multi-head attention that is unique in the architecture. However, little is known for how linguistic properties are processed, represented, and utilized for downstream tasks among hundreds of attention heads inside the pre-trained transformerbased model. For the initial goal of examining the roles of attention heads in handling a set of linguistic features, we conducted a set of experiments with ten probing tasks and three downstream tasks on four pre-trained transformer families (GPT, GPT2, BERT, and ELECTRA). Meaningful insights are shown through the lens of heat map visualization and utilized to propose a relatively simple sentence representation method that takes advantage of most influential attention heads, resulting in additional performance improvements on the downstream tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- ATMAN: Understanding Transformer Predictions Through Memory Efficient Attention ManipulationBjörn Deiseroth, Mayukh Deb, Samuel Weinbach, Manuel Brack 等NeurIPS 2023 · 被引用 45 次
- FIND: Human-in-the-Loop Debugging Deep Text ClassifiersPiyawat Lertvittayakumjorn, Lucia Specia, Francesca ToniEMNLP 2020 · 被引用 33 次
- Artificial Text Detection via Examining the Topology of Attention MapsLaida Kushnareva, Daniil Cherniavskii, Vladislav Mikhailov, Ekaterina Artemova 等EMNLP 2021 · 被引用 27 次
- GPT-D: Inducing Dementia-related Linguistic Anomalies by Deliberate Degradation of Artificial Neural Language ModelsChangye Li, David S. Knopman, Weizhe Xu, Trevor Cohen 等ACL 2022 · 被引用 24 次
- Ultra-High Dimensional Sparse Representations with Binarization for Efficient Text RetrievalKyoungrok Jang, Junmo Kang, Giwon Hong, Sung-Hyon Myaeng 等EMNLP 2021 · 被引用 13 次
它引用的顶会 Paper1
相关 Paper
- Less Mature is More Adaptable for Sentence-level Language ModelingAbhilasha Sancheti, David Dale, Artyom Kozhevnikov, Maha ElbayadACL 2025
- Contributions of Transformer Attention Heads in Multi- and Cross-lingual TasksWeicheng Ma, Kai Zhang, Renze Lou, Lili Wang 等ACL 2021
- Analyzing Individual Neurons in Pre-trained Language ModelsNadir Durrani, Hassan Sajjad, Fahim Dalvi, Yonatan BelinkovEMNLP 2020 · 被引用 5 次
- Map of Encoders - Mapping Sentence Encoders using Quantum Relative EntropyGaifan Zhang, Danushka BollegalaACL 2026
- How does BERT's attention change when you fine-tune? An analysis methodology and a case study in negation scopeYiyun Zhao, Steven BethardACL 2020 · 被引用 35 次
