From SPMRL to NMRL: What Did We Learn (and Unlearn) in a Decade of Parsing Morphologically-Rich Languages (MRLs)?
Reut Tsarfaty, Dan Bareket, Stav Klein, Amit Seker
Abstract
It has been exactly a decade since the first establishment of SPMRL, a research initiative unifying multiple research efforts to address the peculiar challenges of Statistical Parsing for Morphologically-Rich Languages (MRLs). Here we reflect on parsing MRLs in that decade, highlight the solutions and lessons learned for the architectural, modeling and lexical challenges in the pre-neural era, and argue that similar challenges re-emerge in neural architectures for MRLs. We then aim to offer a climax, suggesting that incorporating symbolic ideas proposed in SPMRL terms into nowadays neural architectures has the potential to push NLP for MRLs to a new level. We sketch a strategies for designing Neural Models for MRLs (NMRL), and showcase preliminary support for these strategies via investigating the task of multi-tagging in Hebrew, a morphologically-rich, high-fusion, language.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- AlephBERT: Language Model Pre-training and Evaluation from Sub-Word to Sentence LevelAmit Seker, Elron Bandel, Dan Bareket, Idan Brusilovsky et al.ACL 2022 · 53 citations
- A Large-Scale Study of Machine Translation in Turkic LanguagesJamshidbek Mirzakhalov, Anoop Babu, Duygu Ataman, Sherzod Kariev et al.EMNLP 2021 · 19 citations
- Multi-VALUE: A Framework for Cross-Dialectal English NLPCaleb Ziems, William Barr Held, Jingfeng Yang, Jwala Dhamala et al.ACL 2023 · 15 citations
- You should evaluate your language model on marginal likelihood over tokenisationsKris Cao, Laura RimellEMNLP 2021 · 6 citations
- Evaluating the Morphosyntactic Well-formedness of Generated TextsAdithya Pratapa, Antonios Anastasopoulos, Shruti Rijhwani, Aditi Chaudhary et al.EMNLP 2021 · 6 citations
Related papers
- Beyond Word Boundaries: A Hebrew Coreference Benchmark and an Evaluation Protocol for Morphologically Complex TextRefael Shaked Greenfeld, Reut TsarfatyACL 2026
- Efficient Constituency Parsing by PointingThanh-Tung Nguyen, Xuan-Phi Nguyen, Shafiq R. Joty, Xiaoli LiACL 2020 · 11 citations
- Multilingual Pre-training with Universal Dependency LearningKailai Sun, Zuchao Li, Hai ZhaoNeurIPS 2021 · 11 citations
- Unsupervised Cross-Lingual Part-of-Speech Tagging for Truly Low-Resource ScenariosRamy Eskander, Smaranda Muresan, Michael CollinsEMNLP 2020 · 16 citations
- Exploring morphology-aware tokenization: A case study on Spanish language modelingAlba Táboas García, Piotr Przybyla, Leo WannerEMNLP 2025 · 1 citation
