Probing for Referential Information in Language Models
Ionut-Teodor Sorodoc, Kristina Gulordava, Gemma Boleda
Abstract
Language models keep track of complex linguistic information about the preceding context -including, e.g., syntactic relations in a sentence. We investigate whether they also capture information beneficial for resolving pronominal anaphora in English. We analyze two state of the art models with LSTM and Transformer architectures, respectively, using probe tasks on a coreference annotated corpus. Our hypothesis is that language models will capture grammatical properties of anaphora (such as agreement between a pronoun and its antecedent), but not semantico-referential information (the fact that pronoun and antecedent refer to the same entity). Instead, we find evidence that models capture referential aspects to some extent -though they are still much better at grammar. The Transformer outperforms the LSTM in all analyses, and exhibits in particular better semantico-referential abilities.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Entity Tracking in Language ModelsNajoung Kim, Sebastian SchusterACL 2023 · 9 citations
- 100x Cost & Latency Reduction: Performance Analysis of AI Query Approximation using Lightweight Proxy Models: [Experiments & Analysis]Yeounoh Chung, Rushabh Desai, Jian He, Yu Xiao et al.SIGMOD 2026 · 8 citations
- SocioProbe: What, When, and Where Language Models Learn about SociodemographicsAnne Lauscher, Federico Bianchi, Samuel R. Bowman, Dirk HovyEMNLP 2022 · 6 citations
- Intrinsic Probing through Dimension SelectionLucas Torroba Hennigen, Adina Williams, Ryan CotterellEMNLP 2020 · 3 citations
- Probing Toxic Content in Large Pre-Trained Language ModelsNedjma Ousidhoum, Xinran Zhao, Tianqing Fang, Yangqiu Song et al.ACL 2021
Builds on2
Related papers
- Dependency resolution at the syntax-semantics interface: psycholinguistic and computational insights on control dependenciesIria de-Dios-Flores, Juan Garcia Amboage, Marcos GarcíaACL 2023 · 2 citations
- Coreferential Reasoning Learning for Language RepresentationDeming Ye, Yankai Lin, Jiaju Du, Zhenghao Liu et al.EMNLP 2020 · 164 citations
- Correct-Detect: Balancing Performance and Ambiguity Through the Lens of Coreference Resolution in LLMsAmber Shore, Russell Scheinberg, Ameeta Agrawal, So Young LeeEMNLP 2025
- Seq2seq is All You Need for Coreference ResolutionWenzheng Zhang, Sam Wiseman, Karl StratosEMNLP 2023 · 5 citations
- On the Ability and Limitations of Transformers to Recognize Formal LanguagesSatwik Bhattamishra, Kabir Ahuja, Navin GoyalEMNLP 2020 · 7 citations
