Spying on Your Neighbors: Fine-grained Probing of Contextual Embeddings for Information about Surrounding Words
Josef Klafka, Allyson Ettinger
Abstract
Although models using contextual word embeddings have achieved state-of-the-art results on a host of NLP tasks, little is known about exactly what information these embeddings encode about the context words that they are understood to reflect. To address this question, we introduce a suite of probing tasks that enable fine-grained testing of contextual embeddings for encoding of information about surrounding words. We apply these tasks to examine the popular BERT, ELMo and GPT contextual encoders, and find that each of our tested information types is indeed encoded as contextual information across tokens, often with nearperfect recoverability-but the encoders vary in which features they distribute to which tokens, how nuanced their distributions are, and how robust the encoding of each feature is to distance. We discuss implications of these results for how different types of models break down and prioritize word-level context information when constructing token embeddings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9b0c4939-c1f1-4cbd-806e-78daa65d63c1Cited by top-tier papers9
- Probing for the Usage of Grammatical NumberKarim Lasri, Tiago Pimentel, Alessandro Lenci, Thierry Poibeau et al.ACL 2022 · 72 citations
- Assessing Phrasal Representation and Composition in TransformersLang Yu, Allyson EttingerEMNLP 2020 · 60 citations
- Neural reality of argument structure constructionsBai Li, Zining Zhu, Guillaume Thomas, Frank Rudzicz et al.ACL 2022 · 38 citations
- AdapLeR: Speeding up Inference by Adaptive Length ReductionAli Modarressi, Hosein Mohebbi, Mohammad Taher PilehvarACL 2022 · 34 citations
- Influence Patterns for Explaining Information Flow in BERTKaiji Lu, Zifan Wang, Piotr Mardziel, Anupam DattaNeurIPS 2021 · 22 citations
Related papers
- On the Robustness of Language Encoders against Grammatical ErrorsFan Yin, Quanyu Long, Tao Meng, Kai-Wei ChangACL 2020 · 32 citations
- A Closer Look at How Fine-tuning Changes BERTYichu Zhou, Vivek SrikumarACL 2022 · 84 citations
- Probing Linguistic Features of Sentence-Level Representations in Relation ExtractionChristoph Alt, Aleksandra Gabryszak, Leonhard HennigACL 2020 · 29 citations
- Evaluating the Robustness of Neural Language Models to Input PerturbationsMilad Moradi, Matthias SamwaldEMNLP 2021 · 64 citations
- Bird's Eye: Probing for Linguistic Graph Structures with a Simple Information-Theoretic ApproachYifan Hou, Mrinmaya SachanACL 2021
