Unsupervised Dependency Graph Network
Yikang Shen, Shawn Tan, Alessandro Sordoni, Peng Li, Jie Zhou, Aaron C. Courville
Abstract
Recent work has identified properties of pretrained self-attention models that mirror those of dependency parse structures. In particular, some self-attention heads correspond well to individual dependency types. Inspired by these developments, we propose a new competitive mechanism that encourages these attention heads to model different dependency relations. We introduce a new model, the Unsupervised Dependency Graph Network (UDGN), that can induce dependency structures from raw corpora and the masked language modeling task. Experiment results show that UDGN achieves very strong unsupervised dependency parsing performance without gold POS tags and any other external information. The competitive gated heads show a strong correlation with human-annotated dependency types. Furthermore, the UDGN can also achieve competitive performance on masked language modeling and sentence textual similarity tasks 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 104106ff-b679-4fc8-bf48-1bf10383efd3Cited by top-tier papers3
- Mixture of Attention Heads: Selecting Attention Heads Per TokenXiaofeng Zhang, Yikang Shen, Zeyu Huang, Jie Zhou et al.EMNLP 2022 · 23 citations
- A Systematic Study of Compositional Syntactic Transformer Language ModelsYida Zhao, Hao Xve, Xiang Hu, Kewei TuACL 2025 · 1 citation
- Dependency Transformer Grammars: Integrating Dependency Structures into Transformer Language ModelsYida Zhao, Chao Lou, Kewei TuACL 2024
Builds on8
- Dependency Graph Enhanced Dual-transformer Structure for Aspect-based Sentiment ClassificationHao Tang, Donghong Ji, Chenliang Li, Qiji ZhouACL 2020 · 332 citations
- StructBERT: Incorporating Language Structures into Pre-training for Deep Language UnderstandingWei Wang, Bin Bi, Ming Yan, Chen Wu et al.ICLR 2020 · 297 citations
- A Systematic Assessment of Syntactic Generalization in Neural Language ModelsJennifer Hu, Jon Gauthier, Peng Qian, Ethan Wilcox et al.ACL 2020 · 124 citations
- GATE: Graph Attention Transformer Encoder for Cross-lingual Relation and Event ExtractionWasi Uddin Ahmad, Nanyun Peng, Kai-Wei ChangAAAI 2021 · 113 citations
- Gradient Estimation with Stochastic Softmax TricksMax B. Paulus, Dami Choi, Daniel Tarlow, Andreas Krause et al.NeurIPS 2020 · 104 citations
Related papers
- StructFormer: Joint Unsupervised Induction of Dependency and Constituency Structure from Masked Language ModelingYikang Shen, Yi Tay, Che Zheng, Dara Bahri et al.ACL 2021
- On Eliciting Syntax from Language Models via HashingYiran Wang, Masao UtiyamaEMNLP 2024
- Semi-Supervised Semantic Dependency Parsing Using CRF AutoencodersZixia Jia, Youmi Ma, Jiong Cai, Kewei TuACL 2020 · 10 citations
- Adapting Unsupervised Syntactic Parsing Methodology for Discourse Dependency ParsingLiwen Zhang, Ge Wang, Wenjuan Han, Kewei TuACL 2021
- Universal Decompositional Semantic ParsingElias Stengel-Eskin, Aaron Steven White, Sheng Zhang, Benjamin Van DurmeACL 2020
