Improving Tokenisation by Alternative Treatment of Spaces
Edward Gow-Smith, Harish Tayyar Madabushi, Carolina Scarton, Aline Villavicencio
Abstract
Tokenisation is the first step in almost all NLP tasks, and state-of-the-art transformer-based language models all use subword tokenisation algorithms to process input text. Existing algorithms have problems, often producing tokenisations of limited linguistic validity and representing equivalent strings differently depending on their position within a word. We hypothesise that these problems hinder the ability of transformer-based models to handle complex words, and suggest that these problems are a result of allowing tokens to include spaces. We thus experiment with an alternative tokenisation approach where spaces are always treated as individual tokens. Specifically, we apply this modification to the BPE and Unigram algorithms. We find that our modified algorithms lead to improved performance on downstream NLP tasks that involve handling complex words, whilst having no detrimental effect on performance in general natural language understanding tasks. Intrinsically, we find that our modified algorithms give more morphologically correct tokenisations, in particular when handling prefixes. Given the results of our experiments, we advocate for always treating spaces as individual tokens as an improved tokenisation method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext da943d99-b18f-4df2-ae7d-1ecffa34f991Cited by top-tier papers5
- Tokenization Is More Than CompressionCraig W. Schmidt, Varshini Reddy, Haoran Zhang, Alec Alameddine et al.EMNLP 2024 · 16 citations
- One Tokenizer To Rule Them All: Emergent Language Plasticity via Multilingual TokenizersDiana Abagyan, Alejandro Salamanca, Andrés Felipe Cruz-Salinas, Kris Cao et al.ACL 2026 · 11 citations
- CharBench: Evaluating the Role of Tokenization in Character-Level TasksOmri Uzan, Yuval PinterAAAI 2026 · 3 citations
- Hints on the data for language modeling of synthetic languages with transformersRodolfo Zevallos, Núria BelACL 2023 · 2 citations
- The Foundations of Tokenization: Statistical and Computational ConcernsJuan Luis Gastaldi, John Terilla, Luca Malagutti, Brian DuSell et al.ICLR 2025 · 1 citation
Builds on7
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 541 citations
- StructBERT: Incorporating Language Structures into Pre-training for Deep Language UnderstandingWei Wang, Bin Bi, Ming Yan, Chen Wu et al.ICLR 2020 · 297 citations
- Rare Words: A Major Problem for Contextualized Embeddings and How to Fix it by Attentive MimickingTimo Schick, Hinrich SchützeAAAI 2020 · 106 citations
Related papers
- Tokenisation over Bounded Alphabets is HardVioleta Kastreva, Philip Whittington, Dennis Komm, Tiago PimentelICLR 2026 · 6 citations
- A Partition Cover Approach to TokenizationJia Peng Lim, Shawn Tan, Davin Choo, Hady W. LauwNeurIPS 2025 · 6 citations
- Exploring morphology-aware tokenization: A case study on Spanish language modelingAlba Táboas García, Piotr Przybyla, Leo WannerEMNLP 2025 · 1 citation
- BPE Gets Picky: Efficient Vocabulary Refinement During Tokenizer TrainingPavel Chizhov, Catherine Arnett, Elizaveta Korotkova, Ivan P. YamshchikovEMNLP 2024 · 1 citation
- Superbizarre Is Not Superb: Derivational Morphology Improves BERT's Interpretation of Complex WordsValentin Hofmann, Janet B. Pierrehumbert, Hinrich SchützeACL 2021
