Effects of sub-word segmentation on performance of transformer language models
Jue Hou, Anisia Katinskaia, Anh-Duc Vu, Roman Yangarber
Abstract
Language modeling is a fundamental task in natural language processing, which has been thoroughly explored with various architectures and hyperparameters. However, few studies focus on the effect of sub-word segmentation on the performance of language models (LMs). In this paper, we compare GPT and BERT models trained with the statistical segmentation algorithm BPE vs. two unsupervised algorithms for morphological segmentation-Morfessor and StateMorph. We train the models for several languages-including ones with very rich morphology-and compare their performance with different segmentation algorithms, vocabulary sizes, and model sizes. The results show that training with morphological segmentation allows the LMs to: 1. achieve lower perplexity, 2. converge more efficiently in terms of training time, and 3. achieve equivalent or better evaluation scores on downstream tasks. Lastly, we show 4. that LMs of smaller size using morphological segmentation can perform comparably to models of larger size trained with BPEboth in terms of (1) perplexity and (3) scores on downstream tasks. Points (2) and (4) impact on sustainability of LMs, since they reduce the model cost: size and computation time. While (2) reduces cost only in the training phase, (4) does so also in the inference phase.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 210d978f-501d-4dca-80e3-f4dc3950c2e5Cited by top-tier papers7
- TokSuite: Measuring the Impact of Tokenizer Choice on Language Model BehaviorGül Sena Altıntaş, Malikeh Ehghaghi, Brian Lester, Fengyuan Liu et al.ICML 2026 · 3 citations
- Extracting Linguistic Information from Large Language Models: Syntactic Relations and Derivational KnowledgeTsedeniya Kinfe Temesgen, Marion Di Marco, Alexander FraserEMNLP 2025 · 2 citations
- MYTE: Morphology-Driven Byte Encoding for Better and Fairer Multilingual Language ModelingTomasz Limisiewicz, Terra Blevins, Hila Gonen, Orevaoghene Ahia et al.ACL 2024 · 1 citation
- The Foundations of Tokenization: Statistical and Computational ConcernsJuan Luis Gastaldi, John Terilla, Luca Malagutti, Brian DuSell et al.ICLR 2025 · 1 citation
- Phonemes to the Rescue: Multilingual Tokenization Based on International Phonetic AlphabetMilan Miletic, Julie Kallini, Ekaterina ShutovaACL 2026
Builds on3
- RuCoLA: Russian Corpus of Linguistic AcceptabilityVladislav Mikhailov, Tatiana Shamardina, Max Ryabinin, Alena Pestova et al.EMNLP 2022 · 19 citations
- CompoundPiece: Evaluating and Improving Decompounding Performance of Language ModelsBenjamin Minixhofer, Jonas Pfeiffer, Ivan VulicEMNLP 2023 · 3 citations
- Superbizarre Is Not Superb: Derivational Morphology Improves BERT's Interpretation of Complex WordsValentin Hofmann, Janet B. Pierrehumbert, Hinrich SchützeACL 2021
Related papers
- Exploring morphology-aware tokenization: A case study on Spanish language modelingAlba Táboas García, Piotr Przybyla, Leo WannerEMNLP 2025 · 1 citation
- Lexically Grounded Subword SegmentationJindrich Libovický, Jindrich HelclEMNLP 2024 · 2 citations
- Subword Segmentation in LLMs: Looking at Inflection and ConsistencyMarion Di Marco, Alexander FraserEMNLP 2024
- Hints on the data for language modeling of synthetic languages with transformersRodolfo Zevallos, Núria BelACL 2023 · 2 citations
- BERT-like Models for Slavic Morpheme SegmentationDmitry Morozov, Lizaveta Astapenka, Anna V. Glazkova, Timur Garipov et al.ACL 2025 · 1 citation
