Linguistic Dependencies and Statistical Dependence
Jacob Louis Hoover, Wenyu Du, Alessandro Sordoni, Timothy J. O'Donnell
Abstract
Are pairs of words that tend to occur together also likely to stand in a linguistic dependency? This empirical question is motivated by a long history of literature in cognitive science, psycholinguistics, and NLP. In this work we contribute an extensive analysis of the relationship between linguistic dependencies and statistical dependence between words. Improving on previous work, we introduce the use of large pretrained language models to compute contextualized estimates of the pointwise mutual information between words (CPMI). For multiple models and languages, we extract dependency trees which maximize CPMI, and compare to gold standard linguistic dependencies. Overall, we find that CPMI dependencies achieve an unlabelled undirected attachment score of at most ≈ 0.5. While far above chance, and consistently above a non-contextualized PMI baseline, this score is generally comparable to a simple baseline formed by connecting adjacent words. We analyze which kinds of linguistic dependencies are best captured in CPMI dependencies, and also find marked differences between the estimates of the large pretrained language models, illustrating how their different training schemes affect the type of dependencies they capture.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fc9d41c6-5bb3-41e4-b8f7-ee6982ebdebdCited by top-tier papers3
- Constructions are Revealed in Word DistributionsJoshua Rozner, Leonie Weissweiler, Kyle Mahowald, Cory ShainEMNLP 2025 · 8 citations
- Similarity-weighted Construction of Contextualized Commonsense Knowledge Graphs for Knowledge-intense Argumentation TasksMoritz Plenz, Juri Opitz, Philipp Heinisch, Philipp Cimiano et al.ACL 2023 · 4 citations
- Syntactic Substitutability as Unsupervised Dependency SyntaxJasper Jian, Siva ReddyEMNLP 2023
Builds on4
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Perturbed Masking: Parameter-free Probing for Analyzing and Interpreting BERTZhiyong Wu, Yun Chen, Ben Kao, Qun LiuACL 2020 · 158 citations
- Are Pre-trained Language Models Aware of Phrases? Simple but Strong Baselines for Grammar InductionTaeuk Kim, Jihun Choi, Daniel Edmiston, Sang-goo LeeICLR 2020 · 92 citations
- Exploiting Syntactic Structure for Better Language Modeling: A Syntactic Distance ApproachWenyu Du, Zhouhan Lin, Yikang Shen, Timothy J. O'Donnell et al.ACL 2020 · 15 citations
Related papers
- PMIScore: An Unsupervised Approach to Quantify Dialogue EngagementYongkang Guo, Zhihuan Huang, Yuqing KongWWW 2026
- Neural Methods for Point-wise Dependency EstimationYao-Hung Hubert Tsai, Han Zhao, Makoto Yamada, Louis-Philippe Morency et al.NeurIPS 2020 · 41 citations
- An Information-theoretical Framework for Understanding Out-of-distribution Detection with Pretrained Vision-Language ModelsBo Peng, Jie Lu, Guangquan Zhang, Zhen FangNeurIPS 2025 · 9 citations
- Language models and brains align due to more than next-word prediction and word-level informationGabriele Merlin, Mariya TonevaEMNLP 2024 · 2 citations
- Fine-grained Analysis of Brain-LLM Alignment through Input AttributionMichela Proietti, Roberto Capobianco, Mariya TonevaICML 2026
