Causal Estimation of Tokenisation Bias
Pietro Lesci, Clara Meister, Thomas Hofmann, Andreas Vlachos, Tiago Pimentel
Abstract
Modern language models are typically trained over subword sequences, but ultimately define probabilities over character-strings. Ideally, the choice of the tokeniser-which maps characterstrings to subwords-should not affect the probability assigned to the underlying characterstring; in practice, it does. We define this mismatch as tokenisation bias. In this work, we quantify one particular type of tokenisation bias: the effect of including or not a subword (e.g., ⟨hello⟩) in a tokeniser's vocabulary on the probability a trained model assigns to the corresponding characters (i.e., "hello"). Estimating this effect is challenging because each model is trained with only one tokeniser. We address this by framing tokenisation bias as a causal effect and estimating it using the regression discontinuity design. Specifically, we exploit the fact that tokenisation algorithms rank subwords and add the first K to a tokeniser's vocabulary, where K is an arbitrary cutoff point. As such, we can estimate a causal effect by comparing similar subwords around this cutoff. Experimentally, we find that tokenisation consistently affects models' outputs across scales, vocabularies, and tokenisers. Notably, a subword's presence in a small model's vocabulary may increase its characters' probability by up to 17 times, highlighting tokenisation as a key design choice in language modelling.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 185e9bc2-514e-4d98-9df7-83d28378ff68Cited by top-tier papers6
- Inoculation Prompting: Eliciting traits from LLMs during training can reduce trait expression at test-timeDaniel Tan, Anders Woodruff, Niels Warncke, Arun Jose et al.ICLR 2026 · 16 citations
- Parity-Aware Byte-Pair Encoding: Improving Cross-lingual Fairness in TokenizationNegar Foroutan, Clara Meister, Debjit Paul, Joel Niklaus et al.ACL 2026 · 14 citations
- Token Distillation: Attention-Aware Input Embeddings for New TokensKonstantin Dobler, Desmond Elliott, Gerard de MeloICLR 2026 · 7 citations
- Exploiting Vocabulary Frequency Imbalance in Language Model Pre-trainingWoojin Chung, Jeonghoon KimNeurIPS 2025 · 6 citations
- Using Information Theory to Characterize Prosodic Typology: The Case of Tone, Pitch-Accent and Stress-AccentEthan Wilcox, Cui Ding, Giovanni Acampa, Tiago Pimentel et al.ACL 2025 · 3 citations
Builds on14
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
- Scaling Vision TransformersXiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, Lucas BeyerCVPR 2022 · 767 citations
- Language Model Tokenizers Introduce Unfairness Between LanguagesAleksandar Petrov, Emanuele La Malfa, Philip H. S. Torr, Adel BibiNeurIPS 2023 · 301 citations
- OLMo: Accelerating the Science of Language ModelsDirk Groeneveld, Iz Beltagy, Evan Pete Walsh, Akshita Bhagia et al.ACL 2024 · 52 citations
- An Analysis of Tokenization: Transformers under Markov DataNived Rajaraman, Jiantao Jiao, Kannan RamchandranNeurIPS 2024 · 16 citations
Related papers
- You should evaluate your language model on marginal likelihood over tokenisationsKris Cao, Laura RimellEMNLP 2021 · 6 citations
- How to Compute the Probability of a WordTiago Pimentel, Clara MeisterEMNLP 2024 · 2 citations
- StochasTok: Improving Fine-Grained Subword Understanding in LLMsAnya Sims, Thomas Foster, T. Duy Nguyen-Hien, Klara Kaleb et al.ICLR 2026 · 8 citations
- Exact Byte-Level Probabilities from Tokenized Language Models for FIM-Tasks and Model EnsemblesBuu Phan, Brandon Amos, Itai Gat, Marton Havasi et al.ICLR 2025
- Beyond Text Compression: Evaluating Tokenizers Across ScalesJonas F. Lotz, António Vilarinho Lopes, Stephan Peitz, Hendra Setiawan et al.ACL 2025 · 3 citations
