Exploring morphology-aware tokenization: A case study on Spanish language modeling
Alba Táboas García, Piotr Przybyla, Leo Wanner
Abstract
This paper investigates to what extent the integration of morphological information can improve subword tokenization and thus also language modeling performance. We focus on Spanish, a language with fusional morphology, where subword segmentation can benefit from linguistic structure. Instead of relying on purely data-driven strategies like Byte Pair Encoding (BPE), we explore a linguistically grounded approach: training a tokenizer on morphologically segmented data. To do so, we develop a semi-supervised segmentation model for Spanish, building gold-standard datasets to guide and evaluate it. We then use this tokenizer to pre-train a masked language model and assess its performance on several downstream tasks. Our results show improvements over a baseline with a standard tokenizer, supporting our hypothesis that morphology-aware tokenization offers a viable and principled alternative for improving language modeling.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on4
- A Systematic Assessment of Syntactic Generalization in Neural Language ModelsJennifer Hu, Jon Gauthier, Peng Qian, Ethan Wilcox et al.ACL 2020 · 124 citations
- Mind Your Inflections! Improving NLP for Non-Standard Englishes with Base-Inflection EncodingSamson Tan, Shafiq R. Joty, Lav R. Varshney, Min-Yen KanEMNLP 2020 · 26 citations
- DagoBERT: Generating Derivational Morphology with a Pretrained Language ModelValentin Hofmann, Janet B. Pierrehumbert, Hinrich SchützeEMNLP 2020 · 1 citation
- Superbizarre Is Not Superb: Derivational Morphology Improves BERT's Interpretation of Complex WordsValentin Hofmann, Janet B. Pierrehumbert, Hinrich SchützeACL 2021
Related papers
- Effects of sub-word segmentation on performance of transformer language modelsJue Hou, Anisia Katinskaia, Anh-Duc Vu, Roman YangarberEMNLP 2023 · 2 citations
- KinyaBERT: a Morphology-aware Kinyarwanda Language ModelAntoine Nzeyimana, Andre Niyongabo RubungoACL 2022 · 45 citations
- Improving Tokenisation by Alternative Treatment of SpacesEdward Gow-Smith, Harish Tayyar Madabushi, Carolina Scarton, Aline VillavicencioEMNLP 2022 · 6 citations
- Lexically Grounded Subword SegmentationJindrich Libovický, Jindrich HelclEMNLP 2024 · 2 citations
- TAMS: Translation-Assisted Morphological SegmentationEnora Rice, Ali Marashian, Luke Gessler, Alexis Palmer et al.ACL 2024
