Subword Segmentation in LLMs: Looking at Inflection and Consistency
Marion Di Marco, Alexander Fraser
摘要
The role of subword segmentation in relation to capturing morphological patterns in LLMs is currently not well explored. Ideally, one would train large models like GPT using various segmentations and evaluate how well word meanings are captured. Since this is not computationally feasible, we group words according to their segmentation properties and compare how well a model can solve a linguistic task for these groups. We study two criteria: (i) adherence to morpheme boundaries and (ii) the segmentation consistency of the different inflected forms of a lemma. We select word forms with high and low values for these criteria and carry out experiments on GPT-4o's ability to capture verbal inflection for 10 languages. Our results indicate that in particular the criterion of segmentation consistency can help to predict the model's ability to recognize and generate the lemma from an inflected form, providing evidence that subword segmentation is relevant.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Extracting Linguistic Information from Large Language Models: Syntactic Relations and Derivational KnowledgeTsedeniya Kinfe Temesgen, Marion Di Marco, Alexander FraserEMNLP 2025 · 被引用 2 次
- Vocabulary Shapes Cross-Lingual Variation of Word-Order Learnability in Language ModelsJonas Mayer Martins, Jaap Jumelet, Viola Priesemann, Lisa BeinbornACL 2026
它引用的顶会 Paper4
- KinyaBERT: a Morphology-aware Kinyarwanda Language ModelAntoine Nzeyimana, Andre Niyongabo RubungoACL 2022 · 被引用 45 次
- Counting the Bugs in ChatGPT's Wugs: A Multilingual Investigation into the Morphological Capabilities of a Large Language ModelLeonie Weissweiler, Valentin Hofmann, Anjali Kantharuban, Anna Cai 等EMNLP 2023 · 被引用 10 次
- Effects of sub-word segmentation on performance of transformer language modelsJue Hou, Anisia Katinskaia, Anh-Duc Vu, Roman YangarberEMNLP 2023 · 被引用 2 次
- Superbizarre Is Not Superb: Derivational Morphology Improves BERT's Interpretation of Complex WordsValentin Hofmann, Janet B. Pierrehumbert, Hinrich SchützeACL 2021
相关 Paper
- Exploring morphology-aware tokenization: A case study on Spanish language modelingAlba Táboas García, Piotr Przybyla, Leo WannerEMNLP 2025 · 被引用 1 次
- Massively Multilingual Joint Segmentation and GlossingMichael Ginn, Lindia Tjuatja, Enora Rice, Ali Marashian 等ACL 2026 · 被引用 2 次
- Fairness in Representation for Multilingual NLP: Insights from Controlled Experiments on Conditional Language ModelingAda WanICLR 2022 · 被引用 18 次
- Morphological Inflection: A Reality CheckJordan Kodner, Sarah R. B. Payne, Salam Khalifa, Zoey LiuACL 2023 · 被引用 6 次
- Segment First or Comprehend First? Explore the Limit of Unsupervised Word Segmentation with Large Language ModelsZihong Zhang, Liqi He, Zuchao Li, Lefei Zhang 等ACL 2025 · 被引用 1 次
