Semisupervised Neural Proto-Language Reconstruction
Liang Lu, Peirong Xie, David R. Mortensen
摘要
Existing work implementing comparative reconstruction of ancestral languages (protolanguages) has usually required full supervision. However, historical reconstruction models are only of practical value if they can be trained with a limited amount of labeled data. We propose a semisupervised historical reconstruction task in which the model is trained on only a small amount of labeled data (cognate sets with proto-forms) and a large amount of unlabeled data (cognate sets without proto-forms). We propose a neural architecture for comparative reconstruction (DPD-BiReconstructor) incorporating an essential insight from linguists' comparative method: that reconstructed words should not only be reconstructable from their daughter words, but also deterministically transformable back into their daughter words. We show that this architecture is able to leverage unlabeled cognate sets to outperform strong semisupervised baselines on this novel task 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper3
- The CRINGE Loss: Learning what language not to modelLeonard Adolphs, Tianyu Gao, Jing Xu, Kurt Shuster 等ACL 2023 · 被引用 7 次
- Neural Unsupervised Reconstruction of Protolanguage Word FormsAndre He, Nicholas Tomlin, Dan KleinACL 2023 · 被引用 4 次
- Cognate Transformer for Automated Phonological Reconstruction and Cognate Reflex PredictionV. S. D. S. Mahesh Akavarapu, Arnab BhattacharyaEMNLP 2023 · 被引用 2 次
相关 Paper
- Single-/Multi-Source Cross-Lingual NER via Teacher-Student Learning on Unlabeled Data in Target LanguageQianhui Wu, Zijia Lin, Börje Karlsson, Jianguang Lou 等ACL 2020 · 被引用 59 次
- Verba volant, scripta volant? Don't worry! There are computational solutions for protoword reconstructionLiviu P. Dinu, Ana Sabina Uban, Alina Maria Cristea, Ioan-Bogdan Iordache 等EMNLP 2024
- Semi-supervised Human Pose Estimation in Art-historical ImagesMatthias Springstein, Stefanie Schneider, Christian Althaus, Ralph EwerthACM MM 2022 · 被引用 17 次
- Few-Shot Named Entity Recognition: An Empirical Baseline StudyJiaxin Huang, Chunyuan Li, Krishan Subudhi, Damien Jose 等EMNLP 2021 · 被引用 97 次
- Text Classification Using Label Names Only: A Language Model Self-Training ApproachYu Meng, Yunyi Zhang, Jiaxin Huang, Chenyan Xiong 等EMNLP 2020 · 被引用 203 次
