Are Pre-trained Language Models Aware of Phrases? Simple but Strong Baselines for Grammar Induction
Taeuk Kim, Jihun Choi, Daniel Edmiston, Sang-goo Lee
Abstract
With the recent success and popularity of pre-trained language models (LMs) in natural language processing, there has been a rise in efforts to understand their inner workings. In line with such interest, we propose a novel method that assists us in investigating the extent to which pre-trained LMs capture the syntactic notion of constituency. Our method provides an effective way of extracting constituency trees from the pre-trained LMs without training. In addition, we report intriguing findings in the induced trees, including the fact that pre-trained LMs outperform other approaches in correctly demarcating adverb phrases in sentences.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6f2dc6e9-840a-40d3-bbdc-a8b0d301e4ebCited by top-tier papers23
- Why Do Pretrained Language Models Help in Downstream Tasks? An Analysis of Head and Prompt TuningColin Wei, Sang Michael Xie, Tengyu MaNeurIPS 2021 · 119 citations
- What Do They Capture? - A Structural Analysis of Pre-Trained Language Models for Source CodeYao Wan, Wei Zhao, Hongyu Zhang, Yulei Sui et al.ICSE 2022 · 66 citations
- OA-Mine: Open-World Attribute Mining for E-Commerce Products with Weak SupervisionXinyang Zhang, Chenwei Zhang, Xian Li, Xin Luna Dong et al.WWW 2022 · 36 citations
- Unsupervised Parsing with S-DIORA: Single Tree Encoding for Deep Inside-Outside Recursive AutoencodersAndrew Drozdov, Subendhu Rongali, Yi-Pei Chen, Tim O'Gorman et al.EMNLP 2020 · 27 citations
- Unsupervised Parsing via Constituency TestsSteven Cao, Nikita Kitaev, Dan KleinEMNLP 2020 · 25 citations
Related papers
- Contextual Distortion Reveals Constituency: Masked Language Models are Implicit ParsersJiaxi Li, Wei LuACL 2023 · 3 citations
- Large Language Models Are No Longer Shallow ParsersYuanhe Tian, Fei Xia, Yan SongACL 2024
- AST-Probe: Recovering abstract syntax trees from hidden representations of pre-trained language modelsJosé Antonio Hernández López, Martin Weyssow, Jesús Sánchez Cuadrado, Houari A. SahraouiASE 2022 · 18 citations
- Phrase-aware Unsupervised Constituency ParsingXiaotao Gu, Yikang Shen, Jiaming Shen, Jingbo Shang et al.ACL 2022
- LLM-enhanced Self-training for Cross-domain Constituency ParsingJianling Li, Meishan Zhang, Peiming Guo, Min Zhang et al.EMNLP 2023 · 3 citations
