Word Segmentation as Unsupervised Constituency Parsing
Raquel G. Alhama
摘要
Word identification from continuous input is typically viewed as a segmentation task. Experiments with human adults suggest that familiarity with syntactic structures in their native language also influences word identification in artificial languages; however, the relation between syntactic processing and word identification is yet unclear. This work takes one step forward by exploring a radically different approach of word identification, in which segmentation of a continuous input is viewed as a process isomorphic to unsupervised constituency parsing. Besides formalizing the approach, this study reports simulations of human experiments with DIORA (Drozdov et al., 2020), a neural unsupervised constituency parser. Results show that this model can reproduce human behavior in word identification experiments, suggesting that this is a viable approach to study word identification and its relation to syntactic processing.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
相关 Paper
- Unsupervised Parsing with S-DIORA: Single Tree Encoding for Deep Inside-Outside Recursive AutoencodersAndrew Drozdov, Subendhu Rongali, Yi-Pei Chen, Tim O'Gorman 等EMNLP 2020 · 被引用 27 次
- Improved Latent Tree Induction with Distant Supervision via Span ConstraintsZhiyang Xu, Andrew Drozdov, Jay-Yoon Lee, Tim O'Gorman 等EMNLP 2021 · 被引用 4 次
- Unsupervised Parsing via Constituency TestsSteven Cao, Nikita Kitaev, Dan KleinEMNLP 2020 · 被引用 25 次
- Phrase-aware Unsupervised Constituency ParsingXiaotao Gu, Yikang Shen, Jiaming Shen, Jingbo Shang 等ACL 2022
- Don't Parse, Choose Spans! Continuous and Discontinuous Constituency Parsing via Autoregressive Span SelectionSonglin Yang, Kewei TuACL 2023 · 被引用 1 次
