Confounding Factors in Relating Model Performance to Morphology
Wessel Poelman, Thomas Bauwens, Miryam de Lhoneux
Abstract
The extent to which individual language characteristics influence tokenization and language modeling is an open question. Differences in morphological systems have been suggested as both unimportant and crucial to consider (Cotterell et al., 2018; Gerz et al., 2018a; Park et al., 2021, inter alia). We argue this conflicting evidence is due to confounding factors in experimental setups, making it hard to compare results and draw conclusions. We identify such factors in analyses trying to answer the question of whether, and how, morphology relates to language modeling. Next, we re-assess three hypotheses by Arnett and Bergen (2025) for why modeling agglutinative languages results in higher perplexities than fusional languages: they look at morphological alignment of tokenization, tokenization efficiency, and dataset size. We show that each conclusion includes confounding factors and suggest methodological improvements. Finally, we introduce token bigram metrics as an intrinsic way to predict the difficulty of causal language modeling, and find that they are gradient proxies for morphological complexity that do not require expert annotation. Ultimately, we outline necessities to reliably answer whether, and how, morphology relates to language modeling.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5b3e5ee1-abaa-41e5-9297-a36ea71b9ab0Cited by top-tier papers1
Ask how each one uses itBuilds on6
- Fairness in Representation for Multilingual NLP: Insights from Controlled Experiments on Conditional Language ModelingAda WanICLR 2022 · 18 citations
- Tokenization Is More Than CompressionCraig W. Schmidt, Varshini Reddy, Haoran Zhang, Alec Alameddine et al.EMNLP 2024 · 16 citations
- Tokenization and the Noiseless ChannelVilém Zouhar, Clara Meister, Juan Luis Gastaldi, Li Du et al.ACL 2023 · 10 citations
- Where's the Point? Self-Supervised Multilingual Punctuation-Agnostic Sentence SegmentationBenjamin Minixhofer, Jonas Pfeiffer, Ivan VulicACL 2023 · 8 citations
- MYTE: Morphology-Driven Byte Encoding for Better and Fairer Multilingual Language ModelingTomasz Limisiewicz, Terra Blevins, Hila Gonen, Orevaoghene Ahia et al.ACL 2024 · 1 citation
Related papers
- Hints on the data for language modeling of synthetic languages with transformersRodolfo Zevallos, Núria BelACL 2023 · 2 citations
- Beyond Text Compression: Evaluating Tokenizers Across ScalesJonas F. Lotz, António Vilarinho Lopes, Stephan Peitz, Hendra Setiawan et al.ACL 2025 · 3 citations
- Exploring morphology-aware tokenization: A case study on Spanish language modelingAlba Táboas García, Piotr Przybyla, Leo WannerEMNLP 2025 · 1 citation
- Beyond N-Grams: Rethinking Evaluation Metrics and Strategies for Multilingual Abstractive SummarizationItai Mondshine, Tzuf Paz-Argaman, Reut TsarfatyACL 2025 · 6 citations
- Your Model is Overconfident, and Other Lies We Tell OurselvesTimothee Mickus, Aman Sinha, Raúl VázquezACL 2025
