Exploring morphology-aware tokenization: A case study on Spanish language modeling
Alba Táboas García, Piotr Przybyla, Leo Wanner
摘要
This paper investigates to what extent the integration of morphological information can improve subword tokenization and thus also language modeling performance. We focus on Spanish, a language with fusional morphology, where subword segmentation can benefit from linguistic structure. Instead of relying on purely data-driven strategies like Byte Pair Encoding (BPE), we explore a linguistically grounded approach: training a tokenizer on morphologically segmented data. To do so, we develop a semi-supervised segmentation model for Spanish, building gold-standard datasets to guide and evaluate it. We then use this tokenizer to pre-train a masked language model and assess its performance on several downstream tasks. Our results show improvements over a baseline with a standard tokenizer, supporting our hypothesis that morphology-aware tokenization offers a viable and principled alternative for improving language modeling.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- A Systematic Assessment of Syntactic Generalization in Neural Language ModelsJennifer Hu, Jon Gauthier, Peng Qian, Ethan Wilcox 等ACL 2020 · 被引用 124 次
- Mind Your Inflections! Improving NLP for Non-Standard Englishes with Base-Inflection EncodingSamson Tan, Shafiq R. Joty, Lav R. Varshney, Min-Yen KanEMNLP 2020 · 被引用 26 次
- DagoBERT: Generating Derivational Morphology with a Pretrained Language ModelValentin Hofmann, Janet B. Pierrehumbert, Hinrich SchützeEMNLP 2020 · 被引用 1 次
- Superbizarre Is Not Superb: Derivational Morphology Improves BERT's Interpretation of Complex WordsValentin Hofmann, Janet B. Pierrehumbert, Hinrich SchützeACL 2021
相关 Paper
- Effects of sub-word segmentation on performance of transformer language modelsJue Hou, Anisia Katinskaia, Anh-Duc Vu, Roman YangarberEMNLP 2023 · 被引用 2 次
- KinyaBERT: a Morphology-aware Kinyarwanda Language ModelAntoine Nzeyimana, Andre Niyongabo RubungoACL 2022 · 被引用 45 次
- Improving Tokenisation by Alternative Treatment of SpacesEdward Gow-Smith, Harish Tayyar Madabushi, Carolina Scarton, Aline VillavicencioEMNLP 2022 · 被引用 6 次
- Lexically Grounded Subword SegmentationJindrich Libovický, Jindrich HelclEMNLP 2024 · 被引用 2 次
- TAMS: Translation-Assisted Morphological SegmentationEnora Rice, Ali Marashian, Luke Gessler, Alexis Palmer 等ACL 2024
