Phrase-aware Unsupervised Constituency Parsing
Xiaotao Gu, Yikang Shen, Jiaming Shen, Jingbo Shang, Jiawei Han
摘要
Recent studies have achieved inspiring success in unsupervised grammar induction using masked language modeling (MLM) as the proxy task. Despite their high accuracy in identifying low-level structures, prior arts tend to struggle in capturing high-level structures like clauses, since the MLM task usually only requires information from local context. In this work, we revisit LM-based constituency parsing from a phrase-centered perspective. Inspired by the natural reading process of human readers, we propose to regularize the parser with phrases extracted by an unsupervised phrase tagger to help the LM model quickly manage low-level structures. For a better understanding of high-level structures, we propose a phrase-guided masking strategy for LM to emphasize more on reconstructing nonphrase words. We show that the initial phrase regularization serves as an effective bootstrap, and phrase-guided masking improves the identification of high-level structures. Experiments on the public benchmark with two different backbone models demonstrate the effectiveness and generality of our method.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper5
- Are Pre-trained Language Models Aware of Phrases? Simple but Strong Baselines for Grammar InductionTaeuk Kim, Jihun Choi, Daniel Edmiston, Sang-goo LeeICLR 2020 · 被引用 92 次
- Unsupervised Parsing via Constituency TestsSteven Cao, Nikita Kitaev, Dan KleinEMNLP 2020 · 被引用 25 次
- UCPhrase: Unsupervised Context-aware Quality Phrase TaggingXiaotao Gu, Zihan Wang, Zhenyu Bi, Yu Meng 等KDD 2021 · 被引用 17 次
- Neural Bi-Lexicalized PCFG InductionSonglin Yang, Yanpeng Zhao, Kewei TuACL 2021
- StructFormer: Joint Unsupervised Induction of Dependency and Constituency Structure from Masked Language ModelingYikang Shen, Yi Tay, Che Zheng, Dara Bahri 等ACL 2021
相关 Paper
- Visually Grounded Compound PCFGsYanpeng Zhao, Ivan TitovEMNLP 2020 · 被引用 35 次
- Contextual Distortion Reveals Constituency: Masked Language Models are Implicit ParsersJiaxi Li, Wei LuACL 2023 · 被引用 3 次
- LLM-enhanced Self-training for Cross-domain Constituency ParsingJianling Li, Meishan Zhang, Peiming Guo, Min Zhang 等EMNLP 2023 · 被引用 3 次
- On Eliciting Syntax from Language Models via HashingYiran Wang, Masao UtiyamaEMNLP 2024
- Do Transformers Parse while Predicting the Masked Word?Haoyu Zhao, Abhishek Panigrahi, Rong Ge, Sanjeev AroraEMNLP 2023 · 被引用 5 次
