Lexically Grounded Subword Segmentation
Jindrich Libovický, Jindrich Helcl
摘要
We present three innovations in tokenization and subword segmentation. First, we propose to use unsupervised morphological analysis with Morfessor as pre-tokenization. Second, we present an algebraic method for obtaining subword embeddings grounded in a word embedding space. Based on that, we design a novel subword segmentation algorithm that uses the embeddings, ensuring that the procedure considers lexical meaning. Third, we introduce an efficient segmentation algorithm based on a subword bigram model that can be initialized with the lexically aware segmentation method to avoid using Morfessor and large embedding tables at inference time. We evaluate the proposed approaches using two intrinsic metrics and measure their performance on two downstream tasks: part-of-speech tagging and machine translation. Our experiments show significant improvements in the morphological plausibility of the segmentation when evaluated using segmentation precision on morpheme boundaries and improved Rényi efficiency in 8 languages. Although the proposed tokenization methods do not have a large impact on automatic translation quality, we observe consistent performance gains in the arguably more morphological task of part-of-speech tagging. 1 We use the word morpheme for morphologically motivated subword units. Some theories (Žabokrtský et al., 2022) distinguish morphs as surface realizations of abstract morphemes as the smallest units of meaning. Where appropriate, we follow this distinction for clarity. By morpheme boundaries, we mean boundaries between morphs within a word.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
- Dynamic Programming Encoding for Subword Segmentation in Neural Machine TranslationXuanli He, Gholamreza Haffari, Mohammad NorouziACL 2020 · 被引用 33 次
- Tokenization and the Noiseless ChannelVilém Zouhar, Clara Meister, Juan Luis Gastaldi, Li Du 等ACL 2023 · 被引用 10 次
- Vocabulary Learning via Optimal Transport for Neural Machine TranslationJingjing Xu, Hao Zhou, Chun Gan, Zaixiang Zheng 等ACL 2021
相关 Paper
- Effects of sub-word segmentation on performance of transformer language modelsJue Hou, Anisia Katinskaia, Anh-Duc Vu, Roman YangarberEMNLP 2023 · 被引用 2 次
- Exploring morphology-aware tokenization: A case study on Spanish language modelingAlba Táboas García, Piotr Przybyla, Leo WannerEMNLP 2025 · 被引用 1 次
- T-FREE: Subword Tokenizer-Free Generative LLMs via Sparse Representations for Memory-Efficient EmbeddingsBjörn Deiseroth, Manuel Brack, Patrick Schramowski, Kristian Kersting 等EMNLP 2024 · 被引用 2 次
- Improving Tokenisation by Alternative Treatment of SpacesEdward Gow-Smith, Harish Tayyar Madabushi, Carolina Scarton, Aline VillavicencioEMNLP 2022 · 被引用 6 次
- Retrofitting Large Language Models with Dynamic TokenizationDarius Feher, Ivan Vulic, Benjamin MinixhoferACL 2025
