Hints on the data for language modeling of synthetic languages with transformers
Rodolfo Zevallos, Núria Bel
Abstract
Language Models (LM) are becoming more and more useful for providing representations upon which to train Natural Language Processing applications. However, there is now clear evidence that attention-based transformers require a critical amount of language data to produce good enough LMs. The question we have addressed in this paper is to what extent the critical amount of data varies for languages of different morphological typology, in particular those that have a rich inflectional morphology, and whether the tokenization method to preprocess the data can make a difference. These details can be important for low-resource languages that need to plan the production of datasets. We evaluated intrinsically and extrinsically the differences of five different languages with different pretraining dataset sizes and three different tokenization methods for each. The results confirm that the size of the vocabulary due to morphological characteristics is directly correlated with both the LM perplexity and the performance of two typical downstream tasks such as NER identification and POS Tagging. The experiments also provide new evidence that a canonical tokenizer can reduce perplexity by more than a half for a polysynthetic language like Quechua as well as raising macro-F1 score from 0.8 to more than 0.9 in both downstream tasks with a LM trained with only 6M tokens. 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- The Zeno's Paradox of 'Low-Resource' LanguagesHellina Hailu Nigatu, Atnafu Lambebo Tonja, Benjamin Rosman, Thamar Solorio et al.EMNLP 2024 · 10 citations
- TEMA: Token Embeddings Mapping for Enriching Low-Resource Language ModelsRodolfo Zevallos, Núria Bel, Mireia FarrúsEMNLP 2024
Builds on9
- CamemBERT: a Tasty French Language ModelLouis Martin, Benjamin Muller, Pedro Javier Ortiz Suárez, Yoann Dupont et al.ACL 2020 · 703 citations
- A Systematic Assessment of Syntactic Generalization in Neural Language ModelsJennifer Hu, Jon Gauthier, Peng Qian, Ethan Wilcox et al.ACL 2020 · 124 citations
- How much pretraining data do language models need to learn syntax?Laura Pérez-Mayos, Miguel Ballesteros, Leo WannerEMNLP 2021 · 31 citations
- GottBERT: a pure German Language ModelRaphael Scheible, Johann Frei, Fabian Thomczyk, Henry He et al.EMNLP 2024 · 7 citations
- Learning Which Features Matter: RoBERTa Acquires a Preference for Linguistic Generalizations (Eventually)Alex Warstadt, Yian Zhang, Xiaocheng Li, Haokun Liu et al.EMNLP 2020 · 7 citations
Related papers
- How Good is Your Tokenizer? On the Monolingual Performance of Multilingual Language ModelsPhillip Rust, Jonas Pfeiffer, Ivan Vulic, Sebastian Ruder et al.ACL 2021
- Confounding Factors in Relating Model Performance to MorphologyWessel Poelman, Thomas Bauwens, Miryam de LhoneuxEMNLP 2025
- KinyaBERT: a Morphology-aware Kinyarwanda Language ModelAntoine Nzeyimana, Andre Niyongabo RubungoACL 2022 · 45 citations
- Tokenization and Representation Biases in Multilingual Models on Dialectal NLP TasksVani Kanjirangat, Tanja Samardzic, Ljiljana Dolamic, Fabio RinaldiEMNLP 2025
- Effects of sub-word segmentation on performance of transformer language modelsJue Hou, Anisia Katinskaia, Anh-Duc Vu, Roman YangarberEMNLP 2023 · 2 citations
