Do Transformer Models Show Similar Attention Patterns to Task-Specific Human Gaze?
Oliver Eberle, Stephanie Brandl, Jonas Pilot, Anders Søgaard
Abstract
Learned self-attention functions in state-of-theart NLP models often correlate with human attention. We investigate whether self-attention in large-scale pre-trained language models is as predictive of human eye fixation patterns during task-reading as classical cognitive models of human attention. We compare attention functions across two task-specific reading datasets for sentiment analysis and relation extraction. We find the predictiveness of large-scale pretrained self-attention for human attention depends on 'what is in the tail', e.g., the syntactic nature of rare contexts. Further, we observe that task-specific fine-tuning does not increase the correlation with human task-specific reading. Through an input reduction experiment we give complementary insights on the sparsity and fidelity trade-off, showing that lowerentropy attention vectors are more faithful.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0122ff79-10ad-40eb-a608-ec52b8d08069Cited by top-tier papers4
- Rather a Nurse than a Physician - Contrastive Explanations under InvestigationOliver Eberle, Ilias Chalkidis, Laura Cabello, Stephanie BrandlEMNLP 2023 · 3 citations
- Towards a Mechanistic Interpretation of Multi-Step Reasoning Capabilities of Language ModelsYifan Hou, Jiaoda Li, Yu Fei, Alessandro Stolfo et al.EMNLP 2023 · 2 citations
- Eyes Don't Lie: Subjective Hate Annotation and Detection with GazeÖzge Alaçam, Sanne Hoeken, Sina ZarrießEMNLP 2024 · 1 citation
- Rethinking Spiking Self-Attention Mechanism: Implementing a-XNOR Similarity Calculation in Spiking TransformersYichen Xiao, Shuai Wang, Dehao Zhang, Wenjie Wei et al.CVPR 2025
Builds on4
- Masked Language Modeling and the Distributional Hypothesis: Order Word Matters Pre-training for LittleKoustuv Sinha, Robin Jia, Dieuwke Hupkes, Joelle Pineau et al.EMNLP 2021 · 177 citations
- XAI for Transformers: Better Explanations through Conservative PropagationAmeen Ali, Thomas Schnake, Oliver Eberle, Grégoire Montavon et al.ICML 2022 · 144 citations
- Improving Natural Language Processing Tasks with Human Gaze-Guided Neural AttentionEkta Sood, Simon Tannert, Philipp Müller, Andreas BullingNeurIPS 2020 · 91 citations
- How does BERT's attention change when you fine-tune? An analysis methodology and a case study in negation scopeYiyun Zhao, Steven BethardACL 2020 · 35 citations
Related papers
- Entropy- and Distance-Based Predictors From GPT-2 Attention Patterns Predict Reading Times Over and Above GPT-2 SurprisalByung-Doh Oh, William SchulerEMNLP 2022 · 13 citations
- From Human Reading to NLM Understanding: Evaluating the Role of Eye-Tracking Data in Encoder-Based ModelsLuca Dini, Lucia Domenichelli, Dominique Brunato, Felice Dell'OrlettaACL 2025
- Measuring the Impact of (Psycho-)Linguistic and Readability Features and Their Spill Over Effects on the Prediction of Eye Movement PatternsDaniel Wiechmann, Elma KerzACL 2022 · 17 citations
- Thinking Like a Developer? Comparing the Attention of Humans with Neural Models of CodeMatteo Paltenghi, Michael PradelASE 2021 · 21 citations
- Probing for Reading TimesEleftheria Tsipidi, Samuel Kiegeland, Francesco Ignazio Re, Tianyang Xu et al.ACL 2026
