From Language Models over Tokens to Language Models over Characters
Tim Vieira, Benjamin LeBrun, Mario Giulianelli, Juan Luis Gastaldi, Brian DuSell, John Terilla, Timothy J. O'Donnell, Ryan Cotterell
Abstract
Modern language models are internally-and mathematically-distributions over token strings rather than character strings, posing numerous challenges for programmers building user applications on top of them. For example, if a prompt is specified as a character string, it must be tokenized before passing it to the token-level language model. Thus, the tokenizer and consequent processing are very sensitive to the specification of the prompt (e.g., whether the prompt ends with a space or not). This paper presents algorithms for converting token-level language models to character-level ones. We present both exact and approximate algorithms. In the empirical portion of the paper, we benchmark the practical runtime and approximation quality. Across four publicly available language models, we find that-even with a small computation budget-our method is able to accurately approximate the character-level distribution at reasonably fast speeds, and that a significant improvement in the language model's compression rate (bits/byte) is achieved. https://github.com/genlm/genlm-bytes From Language Models over Tokens to Language Models over Characters If we complete the prompt by taking the most likely next token (greedy completion), we generate the following:
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d4f7b1cc-b5e2-459f-bdd7-f2f48474401dCited by top-tier papers19
- Dynamic Chunking for End-to-End Hierarchical Sequence ModelingSukjun Hwang, Brandon Wang, Albert GuICLR 2026 · 76 citations
- Universal Cross-Tokenizer Distillation via Approximate Likelihood MatchingBenjamin Minixhofer, Ivan Vulic, Edoardo Maria PontiNeurIPS 2025 · 48 citations
- Broken Tokens? Your Language Model can Secretly Handle Non-Canonical TokenizationsBrian Siyuan Zheng, Alisa Liu, Orevaoghene Ahia, Jonathan Hayase et al.NeurIPS 2025 · 19 citations
- Sampling from Your Language Model One Byte at a TimeJonathan Hayase, Alisa Liu, Noah Smith, Sewoong OhICML 2026 · 9 citations
- The Impact of Token Granularity on the Predictive Power of Language Model SurprisalByung-Doh Oh, William SchulerACL 2025 · 7 citations
Builds on7
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- Grammar-Constrained Decoding for Structured NLP Tasks without FinetuningSaibo Geng, Martin Josifoski, Maxime Peyrard, Robert WestEMNLP 2023 · 33 citations
- You should evaluate your language model on marginal likelihood over tokenisationsKris Cao, Laura RimellEMNLP 2021 · 6 citations
- How to Compute the Probability of a WordTiago Pimentel, Clara MeisterEMNLP 2024 · 2 citations
- On the Proper Treatment of Tokenization in PsycholinguisticsMario Giulianelli, Luca Malagutti, Juan Luis Gastaldi, Brian DuSell et al.EMNLP 2024 · 2 citations
Related papers
- Transducing Language ModelsVésteinn Snæbjarnarson, Samuel Kiegeland, Tianyu Liu, Reda Boumasmoud et al.ICLR 2026 · 3 citations
- Proxy Compression for Language ModelingLin Zheng, Li Xinyu, Qian Liu, Xiachong Feng et al.ICML 2026 · 3 citations
- CharBench: Evaluating the Role of Tokenization in Character-Level TasksOmri Uzan, Yuval PinterAAAI 2026 · 3 citations
- Exact Byte-Level Probabilities from Tokenized Language Models for FIM-Tasks and Model EnsemblesBuu Phan, Brandon Amos, Itai Gat, Marton Havasi et al.ICLR 2025
- Language Models over Canonical Byte-Pair EncodingsTim Vieira, Tianyu Liu, Clemente Pasti, Yahya Emara et al.ICML 2025
