How does BERT's attention change when you fine-tune? An analysis methodology and a case study in negation scope
Yiyun Zhao, Steven Bethard
Abstract
Large pretrained language models like BERT, after fine-tuning to a downstream task, have achieved high performance on a variety of NLP problems. Yet explaining their decisions is difficult despite recent work probing their internal representations. We propose a procedure and analysis methods that take a hypothesis of how a transformer-based model might encode a linguistic phenomenon, and test the validity of that hypothesis based on a comparison between knowledge-related downstream tasks with downstream control tasks, and measurement of cross-dataset consistency. We apply this methodology to test BERT and RoBERTa on a hypothesis that some attention heads will consistently attend from a word in negation scope to the negation cue. We find that after fine-tuning BERT and RoBERTa on a negation scope task, the average attention head improves its sensitivity to negation and its attention consistency across negation datasets compared to the pre-trained models. However, only the base models (not the large models) improve compared to a control task, indicating there is evidence for a shallow encoding of negation only in the base models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5955fc8d-1a02-447e-b71d-329bdbf55d1dCited by top-tier papers8
- The Stem Cell Hypothesis: Dilemma behind Multi-Task Learning with Transformer EncodersHan He, Jinho D. ChoiEMNLP 2021 · 111 citations
- Perturbation Augmentation for Fairer NLPRebecca Qian, Candace Ross, Jude Fernandes, Eric Michael Smith et al.EMNLP 2022 · 54 citations
- What's the Best Place for an AI Conference, Vancouver or _______: Why Completing Comparative Questions is DifficultAvishai Zagoury, Einat Minkov, Idan Szpektor, William W. CohenAAAI 2021 · 6 citations
- Human Guided Exploitation of Interpretable Attention Patterns in Summarization and Topic SegmentationRaymond Li, Wen Xiao, Linzi Xing, Lanjun Wang et al.EMNLP 2022 · 4 citations
- Leveraging Affirmative Interpretations from Negation Improves Natural Language UnderstandingMd Mosharaf Hossain, Eduardo BlancoEMNLP 2022 · 4 citations
Related papers
- Intermediate-Task Transfer Learning with Pretrained Language Models: When and Why Does It Work?Yada Pruksachatkun, Jason Phang, Haokun Liu, Phu Mon Htut et al.ACL 2020 · 168 citations
- How Language Models Process NegationZhejian Zhou, Tianyi Zhou, Robin Jia, Jonathan MayICML 2026
- Probing Linguistic Information for Logical Inference in Pre-trained Language ModelsZeming Chen, Qiyue GaoAAAI 2022 · 11 citations
- Compositional and Lexical Semantics in RoBERTa, BERT and DistilBERT: A Case Study on CoQAIeva Staliunaite, Ignacio IacobacciEMNLP 2020 · 2 citations
- When Do You Need Billions of Words of Pretraining Data?Yian Zhang, Alex Warstadt, Xiaocheng Li, Samuel R. BowmanACL 2021
