Language Models over Canonical Byte-Pair Encodings
Tim Vieira, Tianyu Liu, Clemente Pasti, Yahya Emara, Brian DuSell, Benjamin LeBrun, Mario Giulianelli, Juan Luis Gastaldi, Timothy J. O'Donnell, Ryan Cotterell
摘要
Modern language models represent probability distributions over character strings as distributions over (shorter) token strings derived via a deterministic tokenizer, such as byte-pair encoding. While this approach is highly effective at scaling up language models to large corpora, its current incarnations have a concerning property: the model assigns nonzero probability mass to an exponential number of noncanonical token encodings of each character string--these are token strings that decode to valid character strings but are impossible under the deterministic tokenizer (i.e., they will never be seen in any training corpus, no matter how large). This misallocation is both erroneous, as noncanonical strings never appear in training data, and wasteful, diverting probability mass away from plausible outputs. These are avoidable mistakes! In this work, we propose methods to enforce canonicality in token-level language models, ensuring that only canonical token strings are assigned positive probability. We present two approaches: (1) canonicality by conditioning, leveraging test-time inference strategies without additional training, and (2) canonicality by construction, a model parameterization that guarantees canonical outputs but requires training. We demonstrate that fixing canonicality mistakes improves the likelihood of held-out data for several models and corpora.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Sampling from Your Language Model One Byte at a TimeJonathan Hayase, Alisa Liu, Noah Smith, Sewoong OhICML 2026 · 被引用 9 次
- Cross-Tokenizer Likelihood Scoring Algorithms for Language Model DistillationBuu Phan, Ashish Khisti, Karen UllrichICLR 2026 · 被引用 5 次
- When to Ensemble: Identifying Token-Level Points for Stable and Fast LLM EnsemblingHeecheol Yun, Kwangmin Ki, Jung Hyun Lee, Eunho YangICLR 2026 · 被引用 3 次
- Transducing Language ModelsVésteinn Snæbjarnarson, Samuel Kiegeland, Tianyu Liu, Reda Boumasmoud 等ICLR 2026 · 被引用 3 次
- How Persuasive Is Your Context?Tu Nguyen, Kevin Du, Alexander Miserlis Hoyle, Ryan CotterellEMNLP 2025
它引用的顶会 Paper8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley 等ICML 2023 · 被引用 1,822 次
- Learning to summarize with human feedbackNisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel M. Ziegler 等NeurIPS 2020 · 被引用 124 次
- Grammar-Aligned DecodingKanghee Park, Jiayu Wang, Taylor Berg-Kirkpatrick, Nadia Polikarpova 等NeurIPS 2024 · 被引用 73 次
- Probabilistic Inference in Language Models via Twisted Sequential Monte CarloStephen Zhao, Rob Brekelmans, Alireza Makhzani, Roger Baker GrosseICML 2024 · 被引用 61 次
相关 Paper
- Where is the signal in tokenization space?Renato Lui Geh, Honghua Zhang, Kareem Ahmed, Benjie Wang 等EMNLP 2024 · 被引用 1 次
- Beyond Perplexity: UTF-8 Validity in Byte-aware Language ModelsSangwhan Moon, Daisuke Oba, Youmi Ma, Tatsuya Hiraoka 等ICML 2026
- Broken Tokens? Your Language Model can Secretly Handle Non-Canonical TokenizationsBrian Siyuan Zheng, Alisa Liu, Orevaoghene Ahia, Jonathan Hayase 等NeurIPS 2025 · 被引用 19 次
- You should evaluate your language model on marginal likelihood over tokenisationsKris Cao, Laura RimellEMNLP 2021 · 被引用 6 次
- From Language Models over Tokens to Language Models over CharactersTim Vieira, Benjamin LeBrun, Mario Giulianelli, Juan Luis Gastaldi 等ICML 2025
