From Language Models over Tokens to Language Models over Characters
Tim Vieira, Benjamin LeBrun, Mario Giulianelli, Juan Luis Gastaldi, Brian DuSell, John Terilla, Timothy J. O'Donnell, Ryan Cotterell
摘要
Modern language models are internally-and mathematically-distributions over token strings rather than character strings, posing numerous challenges for programmers building user applications on top of them. For example, if a prompt is specified as a character string, it must be tokenized before passing it to the token-level language model. Thus, the tokenizer and consequent processing are very sensitive to the specification of the prompt (e.g., whether the prompt ends with a space or not). This paper presents algorithms for converting token-level language models to character-level ones. We present both exact and approximate algorithms. In the empirical portion of the paper, we benchmark the practical runtime and approximation quality. Across four publicly available language models, we find that-even with a small computation budget-our method is able to accurately approximate the character-level distribution at reasonably fast speeds, and that a significant improvement in the language model's compression rate (bits/byte) is achieved. https://github.com/genlm/genlm-bytes From Language Models over Tokens to Language Models over Characters If we complete the prompt by taking the most likely next token (greedy completion), we generate the following:
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- Dynamic Chunking for End-to-End Hierarchical Sequence ModelingSukjun Hwang, Brandon Wang, Albert GuICLR 2026 · 被引用 76 次
- Universal Cross-Tokenizer Distillation via Approximate Likelihood MatchingBenjamin Minixhofer, Ivan Vulic, Edoardo Maria PontiNeurIPS 2025 · 被引用 48 次
- Broken Tokens? Your Language Model can Secretly Handle Non-Canonical TokenizationsBrian Siyuan Zheng, Alisa Liu, Orevaoghene Ahia, Jonathan Hayase 等NeurIPS 2025 · 被引用 19 次
- Sampling from Your Language Model One Byte at a TimeJonathan Hayase, Alisa Liu, Noah Smith, Sewoong OhICML 2026 · 被引用 9 次
- The Impact of Token Granularity on the Predictive Power of Language Model SurprisalByung-Doh Oh, William SchulerACL 2025 · 被引用 7 次
它引用的顶会 Paper7
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- Grammar-Constrained Decoding for Structured NLP Tasks without FinetuningSaibo Geng, Martin Josifoski, Maxime Peyrard, Robert WestEMNLP 2023 · 被引用 33 次
- You should evaluate your language model on marginal likelihood over tokenisationsKris Cao, Laura RimellEMNLP 2021 · 被引用 6 次
- How to Compute the Probability of a WordTiago Pimentel, Clara MeisterEMNLP 2024 · 被引用 2 次
- On the Proper Treatment of Tokenization in PsycholinguisticsMario Giulianelli, Luca Malagutti, Juan Luis Gastaldi, Brian DuSell 等EMNLP 2024 · 被引用 2 次
相关 Paper
- Transducing Language ModelsVésteinn Snæbjarnarson, Samuel Kiegeland, Tianyu Liu, Reda Boumasmoud 等ICLR 2026 · 被引用 3 次
- Proxy Compression for Language ModelingLin Zheng, Li Xinyu, Qian Liu, Xiachong Feng 等ICML 2026 · 被引用 3 次
- CharBench: Evaluating the Role of Tokenization in Character-Level TasksOmri Uzan, Yuval PinterAAAI 2026 · 被引用 3 次
- Exact Byte-Level Probabilities from Tokenized Language Models for FIM-Tasks and Model EnsemblesBuu Phan, Brandon Amos, Itai Gat, Marton Havasi 等ICLR 2025
- Language Models over Canonical Byte-Pair EncodingsTim Vieira, Tianyu Liu, Clemente Pasti, Yahya Emara 等ICML 2025
