Unsupervised Dependency Graph Network
Yikang Shen, Shawn Tan, Alessandro Sordoni, Peng Li, Jie Zhou, Aaron C. Courville
摘要
Recent work has identified properties of pretrained self-attention models that mirror those of dependency parse structures. In particular, some self-attention heads correspond well to individual dependency types. Inspired by these developments, we propose a new competitive mechanism that encourages these attention heads to model different dependency relations. We introduce a new model, the Unsupervised Dependency Graph Network (UDGN), that can induce dependency structures from raw corpora and the masked language modeling task. Experiment results show that UDGN achieves very strong unsupervised dependency parsing performance without gold POS tags and any other external information. The competitive gated heads show a strong correlation with human-annotated dependency types. Furthermore, the UDGN can also achieve competitive performance on masked language modeling and sentence textual similarity tasks 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Mixture of Attention Heads: Selecting Attention Heads Per TokenXiaofeng Zhang, Yikang Shen, Zeyu Huang, Jie Zhou 等EMNLP 2022 · 被引用 23 次
- A Systematic Study of Compositional Syntactic Transformer Language ModelsYida Zhao, Hao Xve, Xiang Hu, Kewei TuACL 2025 · 被引用 1 次
- Dependency Transformer Grammars: Integrating Dependency Structures into Transformer Language ModelsYida Zhao, Chao Lou, Kewei TuACL 2024
它引用的顶会 Paper8
- Dependency Graph Enhanced Dual-transformer Structure for Aspect-based Sentiment ClassificationHao Tang, Donghong Ji, Chenliang Li, Qiji ZhouACL 2020 · 被引用 332 次
- StructBERT: Incorporating Language Structures into Pre-training for Deep Language UnderstandingWei Wang, Bin Bi, Ming Yan, Chen Wu 等ICLR 2020 · 被引用 297 次
- A Systematic Assessment of Syntactic Generalization in Neural Language ModelsJennifer Hu, Jon Gauthier, Peng Qian, Ethan Wilcox 等ACL 2020 · 被引用 124 次
- GATE: Graph Attention Transformer Encoder for Cross-lingual Relation and Event ExtractionWasi Uddin Ahmad, Nanyun Peng, Kai-Wei ChangAAAI 2021 · 被引用 113 次
- Gradient Estimation with Stochastic Softmax TricksMax B. Paulus, Dami Choi, Daniel Tarlow, Andreas Krause 等NeurIPS 2020 · 被引用 104 次
相关 Paper
- StructFormer: Joint Unsupervised Induction of Dependency and Constituency Structure from Masked Language ModelingYikang Shen, Yi Tay, Che Zheng, Dara Bahri 等ACL 2021
- On Eliciting Syntax from Language Models via HashingYiran Wang, Masao UtiyamaEMNLP 2024
- Semi-Supervised Semantic Dependency Parsing Using CRF AutoencodersZixia Jia, Youmi Ma, Jiong Cai, Kewei TuACL 2020 · 被引用 10 次
- Adapting Unsupervised Syntactic Parsing Methodology for Discourse Dependency ParsingLiwen Zhang, Ge Wang, Wenjuan Han, Kewei TuACL 2021
- Universal Decompositional Semantic ParsingElias Stengel-Eskin, Aaron Steven White, Sheng Zhang, Benjamin Van DurmeACL 2020
