Beyond Perplexity: UTF-8 Validity in Byte-aware Language Models
Sangwhan Moon, Daisuke Oba, Youmi Ma, Tatsuya Hiraoka, Naoaki Okazaki
Abstract
Byte-level tokenization enables language models to handle any Unicode input, but models can generate invalid UTF-8 sequences when encountering rare or unseen characters. We investigate the relationship between training scale and UTF-8 generation reliability with a 355M parameter model trained on 80B tokens from a balanced multilingual corpus of English, Japanese, Korean, and Chinese. We introduce multiple evaluation protocols that isolate UTF-8 structural validity from language modeling. UTF-8 validity convergence lags perplexity by roughly a factor of two: perplexity stabilizes after 2.1B tokens, but UTF-8 validity requires 4.2B tokens. In context-free generation, common characters achieve higher structural validity than rare characters, with bytelength exposure emerging as an additional axis of difficulty alongside frequency. Our experiments show that reliable UTF-8 generation is a distinct capability requiring evaluation beyond perplexity. github.com/cynthia/bytecanary
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 738931fd-a1d2-4209-bc76-6e469e6db671Builds on7
- Neural Machine Translation with Byte-Level SubwordsChanghan Wang, Kyunghyun Cho, Jiatao GuAAAI 2020 · 213 citations
- Byte Latent Transformer: Patches Scale Better Than TokensArtidoro Pagnoni, Ramakanth Pasunuru, Pedro Rodríguez, John Nguyen et al.ACL 2025 · 116 citations
- Glitch Tokens in Large Language Models: Categorization Taxonomy and Effective DetectionYuxi Li, Yi Liu, Gelei Deng, Ying Zhang et al.FSE 2024 · 12 citations
- Fishing for Magikarp: Automatically Detecting Under-trained Tokens in Large Language ModelsSander Land, Max BartoloEMNLP 2024 · 4 citations
- MYTE: Morphology-Driven Byte Encoding for Better and Fairer Multilingual Language ModelingTomasz Limisiewicz, Terra Blevins, Hila Gonen, Orevaoghene Ahia et al.ACL 2024 · 1 citation
Related papers
- Language Models over Canonical Byte-Pair EncodingsTim Vieira, Tianyu Liu, Clemente Pasti, Yahya Emara et al.ICML 2025
- SpaceByte: Towards Deleting Tokenization from Large Language ModelingKevin SlagleNeurIPS 2024 · 34 citations
- Sampling from Your Language Model One Byte at a TimeJonathan Hayase, Alisa Liu, Noah Smith, Sewoong OhICML 2026 · 9 citations
- Beyond Text Compression: Evaluating Tokenizers Across ScalesJonas F. Lotz, António Vilarinho Lopes, Stephan Peitz, Hendra Setiawan et al.ACL 2025 · 3 citations
- Language Model Tokenizers Introduce Unfairness Between LanguagesAleksandar Petrov, Emanuele La Malfa, Philip H. S. Torr, Adel BibiNeurIPS 2023 · 301 citations
