Weakly Supervised Word Segmentation for Computational Language Documentation
Shu Okabe, Laurent Besacier, François Yvon
摘要
Word and morpheme segmentation are fundamental steps of language documentation as they allow to discover lexical units in a language for which the lexicon is unknown. However, in most language documentation scenarios, linguists do not start from a blank page: they may already have a pre-existing dictionary or have initiated manual segmentation of a small part of their data. This paper studies how such a weak supervision can be taken advantage of in Bayesian non-parametric models of segmentation. Our experiments on two very low resource languages (Mboshi and Japhug), whose documentation is still in progress, show that weak supervision can be beneficial to the segmentation quality. In addition, we investigate an incremental learning scenario where manual segmentations are provided in a sequential manner. This work opens the way for interactive annotation tools for documentary linguists.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper1
相关 Paper
- To POS Tag or Not to POS Tag: The Impact of POS Tags on Morphological Learning in Low-Resource SettingsSarah R. Moeller, Ling Liu, Mans HuldenACL 2021
- Tackling the Low-resource Challenge for Canonical SegmentationManuel Mager, Özlem Çetinoglu, Katharina KannEMNLP 2020
- TAMS: Translation-Assisted Morphological SegmentationEnora Rice, Ali Marashian, Luke Gessler, Alexis Palmer 等ACL 2024
- Weakly Supervised POS Taggers Perform Poorly on Truly Low-Resource LanguagesKatharina Kann, Ophélie Lacroix, Anders SøgaardAAAI 2020 · 被引用 21 次
- Massively Multilingual Joint Segmentation and GlossingMichael Ginn, Lindia Tjuatja, Enora Rice, Ali Marashian 等ACL 2026 · 被引用 2 次
