Beyond Perplexity: UTF-8 Validity in Byte-aware Language Models
Sangwhan Moon, Daisuke Oba, Youmi Ma, Tatsuya Hiraoka, Naoaki Okazaki
摘要
Byte-level tokenization enables language models to handle any Unicode input, but models can generate invalid UTF-8 sequences when encountering rare or unseen characters. We investigate the relationship between training scale and UTF-8 generation reliability with a 355M parameter model trained on 80B tokens from a balanced multilingual corpus of English, Japanese, Korean, and Chinese. We introduce multiple evaluation protocols that isolate UTF-8 structural validity from language modeling. UTF-8 validity convergence lags perplexity by roughly a factor of two: perplexity stabilizes after 2.1B tokens, but UTF-8 validity requires 4.2B tokens. In context-free generation, common characters achieve higher structural validity than rare characters, with bytelength exposure emerging as an additional axis of difficulty alongside frequency. Our experiments show that reliable UTF-8 generation is a distinct capability requiring evaluation beyond perplexity. github.com/cynthia/bytecanary
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- Neural Machine Translation with Byte-Level SubwordsChanghan Wang, Kyunghyun Cho, Jiatao GuAAAI 2020 · 被引用 213 次
- Byte Latent Transformer: Patches Scale Better Than TokensArtidoro Pagnoni, Ramakanth Pasunuru, Pedro Rodríguez, John Nguyen 等ACL 2025 · 被引用 116 次
- Glitch Tokens in Large Language Models: Categorization Taxonomy and Effective DetectionYuxi Li, Yi Liu, Gelei Deng, Ying Zhang 等FSE 2024 · 被引用 12 次
- Fishing for Magikarp: Automatically Detecting Under-trained Tokens in Large Language ModelsSander Land, Max BartoloEMNLP 2024 · 被引用 4 次
- MYTE: Morphology-Driven Byte Encoding for Better and Fairer Multilingual Language ModelingTomasz Limisiewicz, Terra Blevins, Hila Gonen, Orevaoghene Ahia 等ACL 2024 · 被引用 1 次
相关 Paper
- Language Models over Canonical Byte-Pair EncodingsTim Vieira, Tianyu Liu, Clemente Pasti, Yahya Emara 等ICML 2025
- SpaceByte: Towards Deleting Tokenization from Large Language ModelingKevin SlagleNeurIPS 2024 · 被引用 34 次
- Sampling from Your Language Model One Byte at a TimeJonathan Hayase, Alisa Liu, Noah Smith, Sewoong OhICML 2026 · 被引用 9 次
- Beyond Text Compression: Evaluating Tokenizers Across ScalesJonas F. Lotz, António Vilarinho Lopes, Stephan Peitz, Hendra Setiawan 等ACL 2025 · 被引用 3 次
- Language Model Tokenizers Introduce Unfairness Between LanguagesAleksandar Petrov, Emanuele La Malfa, Philip H. S. Torr, Adel BibiNeurIPS 2023 · 被引用 301 次
