Broken Tokens? Your Language Model can Secretly Handle Non-Canonical Tokenizations
Brian Siyuan Zheng, Alisa Liu, Orevaoghene Ahia, Jonathan Hayase, Yejin Choi, Noah A. Smith
Abstract
Modern tokenizers employ deterministic algorithms to map text into a single "canonical" token sequence, yet the same string can be encoded as many noncanonical tokenizations using the tokenizer vocabulary. In this work, we investigate the robustness of LMs to text encoded with non-canonical tokenizations entirely unseen during training. Surprisingly, when evaluated across 20 benchmarks, we find that instruction-tuned models retain up to 93.4% of their original performance when given a randomly sampled tokenization, and 90.8% with character-level tokenization. We see that overall stronger models tend to be more robust, and robustness diminishes as the tokenization departs farther from the canonical form. Motivated by these results, we then identify settings where non-canonical tokenization schemes can improve performance, finding that character-level segmentation improves string manipulation and code understanding tasks by up to +14%, and right-aligned digit grouping enhances large-number arithmetic by +33%. Finally, we investigate the source of this robustness, finding that it arises in the instructiontuning phase. We show that while both base and post-trained models grasp the semantics of non-canonical tokenizations (perceiving them as containing misspellings), base models try to mimic the imagined mistakes and degenerate into nonsensical output, while post-trained models are committed to fluent responses. Overall, our findings suggest that models are less tied to their tokenizer than previously believed, and demonstrate the promise of intervening on tokenization at inference time to boost performance. 1 QWEN-2.5-7B-INSTRUCT LLAMA-3.1-8B-INSTRUCT OLMO-2-7B-INSTRUCT Benchmark Canon Rand ∆ Char ∆ Canon Rand ∆ Char ∆ Canon Rand ∆ Char ∆ Multiple choice (MC)
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Proxy Compression for Language ModelingLin Zheng, Li Xinyu, Qian Liu, Xiachong Feng et al.ICML 2026 · 3 citations
- TokDrift: When LLM Speaks in Subwords but Code Speaks in GrammarYinxi Li, Yuntian Deng, Pengyu NieACL 2026 · 2 citations
- Synthetic Object Compositions for Scalable and Accurate Learning in Detection, Segmentation, and GroundingWeikai Huang, Jieyu Zhang, Taoyang jia, Chenhao Zheng et al.CVPR 2026 · 1 citation
- How Tokenization Limits Phonological Knowledge Representation in Language Models and How to Improve ThemDisen Liao, Freda ShiACL 2026 · 1 citation
- so much depends / upon / a whitespace: Why Whitespace Matters for Poets and LLMsSriharsh Bhyravajjula, Melanie Walsh, Anna Preus, Maria AntoniakEMNLP 2025 · 1 citation
Builds on22
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- AlpacaFarm: A Simulation Framework for Methods that Learn from Human FeedbackYann Dubois, Chen Xuechen Li, Rohan Taori, Tianyi Zhang et al.NeurIPS 2023 · 948 citations
- Charformer: Fast Character Transformers via Gradient-based Subword TokenizationYi Tay, Vinh Q. Tran, Sebastian Ruder, Jai Prakash Gupta et al.ICLR 2022 · 198 citations
Related papers
- Where is the signal in tokenization space?Renato Lui Geh, Honghua Zhang, Kareem Ahmed, Benjie Wang et al.EMNLP 2024 · 1 citation
- Understanding the Ability of LLMs to Handle Character-Level PerturbationAnyuan Zhuo, Xuefei Ning, Ningyuan Li, Jingyi Zhu et al.ICML 2026
- StochasTok: Improving Fine-Grained Subword Understanding in LLMsAnya Sims, Thomas Foster, T. Duy Nguyen-Hien, Klara Kaleb et al.ICLR 2026 · 8 citations
- TokSuite: Measuring the Impact of Tokenizer Choice on Language Model BehaviorGül Sena Altıntaş, Malikeh Ehghaghi, Brian Lester, Fengyuan Liu et al.ICML 2026 · 3 citations
- Language Models over Canonical Byte-Pair EncodingsTim Vieira, Tianyu Liu, Clemente Pasti, Yahya Emara et al.ICML 2025
