From SPMRL to NMRL: What Did We Learn (and Unlearn) in a Decade of Parsing Morphologically-Rich Languages (MRLs)?
Reut Tsarfaty, Dan Bareket, Stav Klein, Amit Seker
摘要
It has been exactly a decade since the first establishment of SPMRL, a research initiative unifying multiple research efforts to address the peculiar challenges of Statistical Parsing for Morphologically-Rich Languages (MRLs). Here we reflect on parsing MRLs in that decade, highlight the solutions and lessons learned for the architectural, modeling and lexical challenges in the pre-neural era, and argue that similar challenges re-emerge in neural architectures for MRLs. We then aim to offer a climax, suggesting that incorporating symbolic ideas proposed in SPMRL terms into nowadays neural architectures has the potential to push NLP for MRLs to a new level. We sketch a strategies for designing Neural Models for MRLs (NMRL), and showcase preliminary support for these strategies via investigating the task of multi-tagging in Hebrew, a morphologically-rich, high-fusion, language.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- AlephBERT: Language Model Pre-training and Evaluation from Sub-Word to Sentence LevelAmit Seker, Elron Bandel, Dan Bareket, Idan Brusilovsky 等ACL 2022 · 被引用 53 次
- A Large-Scale Study of Machine Translation in Turkic LanguagesJamshidbek Mirzakhalov, Anoop Babu, Duygu Ataman, Sherzod Kariev 等EMNLP 2021 · 被引用 19 次
- Multi-VALUE: A Framework for Cross-Dialectal English NLPCaleb Ziems, William Barr Held, Jingfeng Yang, Jwala Dhamala 等ACL 2023 · 被引用 15 次
- You should evaluate your language model on marginal likelihood over tokenisationsKris Cao, Laura RimellEMNLP 2021 · 被引用 6 次
- Evaluating the Morphosyntactic Well-formedness of Generated TextsAdithya Pratapa, Antonios Anastasopoulos, Shruti Rijhwani, Aditi Chaudhary 等EMNLP 2021 · 被引用 6 次
相关 Paper
- Beyond Word Boundaries: A Hebrew Coreference Benchmark and an Evaluation Protocol for Morphologically Complex TextRefael Shaked Greenfeld, Reut TsarfatyACL 2026
- Efficient Constituency Parsing by PointingThanh-Tung Nguyen, Xuan-Phi Nguyen, Shafiq R. Joty, Xiaoli LiACL 2020 · 被引用 11 次
- Multilingual Pre-training with Universal Dependency LearningKailai Sun, Zuchao Li, Hai ZhaoNeurIPS 2021 · 被引用 11 次
- Unsupervised Cross-Lingual Part-of-Speech Tagging for Truly Low-Resource ScenariosRamy Eskander, Smaranda Muresan, Michael CollinsEMNLP 2020 · 被引用 16 次
- Exploring morphology-aware tokenization: A case study on Spanish language modelingAlba Táboas García, Piotr Przybyla, Leo WannerEMNLP 2025 · 被引用 1 次
