Do Transformer Models Show Similar Attention Patterns to Task-Specific Human Gaze?
Oliver Eberle, Stephanie Brandl, Jonas Pilot, Anders Søgaard
摘要
Learned self-attention functions in state-of-theart NLP models often correlate with human attention. We investigate whether self-attention in large-scale pre-trained language models is as predictive of human eye fixation patterns during task-reading as classical cognitive models of human attention. We compare attention functions across two task-specific reading datasets for sentiment analysis and relation extraction. We find the predictiveness of large-scale pretrained self-attention for human attention depends on 'what is in the tail', e.g., the syntactic nature of rare contexts. Further, we observe that task-specific fine-tuning does not increase the correlation with human task-specific reading. Through an input reduction experiment we give complementary insights on the sparsity and fidelity trade-off, showing that lowerentropy attention vectors are more faithful.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Rather a Nurse than a Physician - Contrastive Explanations under InvestigationOliver Eberle, Ilias Chalkidis, Laura Cabello, Stephanie BrandlEMNLP 2023 · 被引用 3 次
- Towards a Mechanistic Interpretation of Multi-Step Reasoning Capabilities of Language ModelsYifan Hou, Jiaoda Li, Yu Fei, Alessandro Stolfo 等EMNLP 2023 · 被引用 2 次
- Eyes Don't Lie: Subjective Hate Annotation and Detection with GazeÖzge Alaçam, Sanne Hoeken, Sina ZarrießEMNLP 2024 · 被引用 1 次
- Rethinking Spiking Self-Attention Mechanism: Implementing a-XNOR Similarity Calculation in Spiking TransformersYichen Xiao, Shuai Wang, Dehao Zhang, Wenjie Wei 等CVPR 2025
它引用的顶会 Paper4
- Masked Language Modeling and the Distributional Hypothesis: Order Word Matters Pre-training for LittleKoustuv Sinha, Robin Jia, Dieuwke Hupkes, Joelle Pineau 等EMNLP 2021 · 被引用 177 次
- XAI for Transformers: Better Explanations through Conservative PropagationAmeen Ali, Thomas Schnake, Oliver Eberle, Grégoire Montavon 等ICML 2022 · 被引用 144 次
- Improving Natural Language Processing Tasks with Human Gaze-Guided Neural AttentionEkta Sood, Simon Tannert, Philipp Müller, Andreas BullingNeurIPS 2020 · 被引用 91 次
- How does BERT's attention change when you fine-tune? An analysis methodology and a case study in negation scopeYiyun Zhao, Steven BethardACL 2020 · 被引用 35 次
相关 Paper
- Entropy- and Distance-Based Predictors From GPT-2 Attention Patterns Predict Reading Times Over and Above GPT-2 SurprisalByung-Doh Oh, William SchulerEMNLP 2022 · 被引用 13 次
- From Human Reading to NLM Understanding: Evaluating the Role of Eye-Tracking Data in Encoder-Based ModelsLuca Dini, Lucia Domenichelli, Dominique Brunato, Felice Dell'OrlettaACL 2025
- Measuring the Impact of (Psycho-)Linguistic and Readability Features and Their Spill Over Effects on the Prediction of Eye Movement PatternsDaniel Wiechmann, Elma KerzACL 2022 · 被引用 17 次
- Thinking Like a Developer? Comparing the Attention of Humans with Neural Models of CodeMatteo Paltenghi, Michael PradelASE 2021 · 被引用 21 次
- Probing for Reading TimesEleftheria Tsipidi, Samuel Kiegeland, Francesco Ignazio Re, Tianyang Xu 等ACL 2026
