Tokenization and the Noiseless Channel
Vilém Zouhar, Clara Meister, Juan Luis Gastaldi, Li Du, Mrinmaya Sachan, Ryan Cotterell
Abstract
Subword tokenization is a key part of many NLP pipelines. However, little is known about why some tokenizer and hyperparameter combinations lead to better downstream model performance than others. We propose that good tokenizers lead to efficient channel usage, where the channel is the means by which some input is conveyed to the model and efficiency can be quantified in information-theoretic terms as the ratio of the Shannon entropy to the maximum possible entropy of the token distribution. Yet, an optimal encoding according to Shannon entropy assigns extremely long codes to low-frequency tokens and very short codes to high-frequency tokens. Defining efficiency in terms of Rényi entropy, on the other hand, penalizes distributions with either very high or very low-frequency tokens. In machine translation, we find that across multiple tokenizers, the Rényi entropy with α = 2.5 has a very strong correlation with BLEU: 0.78 in comparison to just -0.32 for compressed length.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext eeff8097-0781-4d02-9935-ae00372fef33Cited by top-tier papers25
- Getting the most out of your tokenizer for pre-training and domain adaptationGautier Dagan, Gabriel Synnaeve, Baptiste RozièreICML 2024 · 68 citations
- An Analysis of Tokenization: Transformers under Markov DataNived Rajaraman, Jiantao Jiao, Kannan RamchandranNeurIPS 2024 · 16 citations
- Tokenization Is More Than CompressionCraig W. Schmidt, Varshini Reddy, Haoran Zhang, Alec Alameddine et al.EMNLP 2024 · 16 citations
- Parity-Aware Byte-Pair Encoding: Improving Cross-lingual Fairness in TokenizationNegar Foroutan, Clara Meister, Debjit Paul, Joel Niklaus et al.ACL 2026 · 14 citations
- Anything Goes? A Crosslinguistic Study of (Im)possible Language Learning in LMsXiulin Yang, Tatsuya Aoyama, Yuekun Yao, Ethan WilcoxACL 2025 · 9 citations
Builds on4
- BLEURT: Learning Robust Metrics for Text GenerationThibault Sellam, Dipanjan Das, Ankur P. ParikhACL 2020 · 40 citations
- BPE-Dropout: Simple and Effective Subword RegularizationIvan Provilkov, Dmitrii Emelianenko, Elena VoitaACL 2020 · 17 citations
- CCAligned: A Massive Collection of Cross-Lingual Web-Document PairsAhmed El-Kishky, Vishrav Chaudhary, Francisco Guzmán, Philipp KoehnEMNLP 2020 · 6 citations
- COMET: A Neural Framework for MT EvaluationRicardo Rei, Craig Stewart, Ana C. Farinha, Alon LavieEMNLP 2020 · 6 citations
Related papers
- Beyond Text Compression: Evaluating Tokenizers Across ScalesJonas F. Lotz, António Vilarinho Lopes, Stephan Peitz, Hendra Setiawan et al.ACL 2025 · 3 citations
- Unsupervised Tokenization LearningAnton Kolonin, Vignav RameshEMNLP 2022 · 3 citations
- Vocabulary Learning via Optimal Transport for Neural Machine TranslationJingjing Xu, Hao Zhou, Chun Gan, Zaixiang Zheng et al.ACL 2021
- Pre-trained Models Perform the Best When Token Distributions Follow Zipf's LawYanjin He, Qingkai Zeng, Meng JiangEMNLP 2025 · 1 citation
- BPE Gets Picky: Efficient Vocabulary Refinement During Tokenizer TrainingPavel Chizhov, Catherine Arnett, Elizaveta Korotkova, Ivan P. YamshchikovEMNLP 2024 · 1 citation
