How to Compute the Probability of a Word
Tiago Pimentel, Clara Meister
Abstract
Language models (LMs) estimate a probability distribution over strings in a natural language; these distributions are crucial for computing perplexity and surprisal in linguistics research. While we are usually concerned with measuring these values for words, most LMs operate over subwords. Despite seemingly straightforward, accurately computing probabilities over one unit given probabilities over the other requires care. Indeed, we show here that many recent linguistic studies have been incorrectly computing these values. This paper derives the correct methods for computing word probabilities, highlighting issues when relying on language models that use beginning-of-word (bow)-marking tokenisers, e.g., the GPT family. Empirically, we show that correcting the widespread bug in probability computations affects measured outcomes in sentence comprehension and lexical optimisation analyses. tpimentelms/probability-of-a-word pip install wordsprobability
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 558ffa2f-c2cd-4fdb-8081-38cdeb6dbf4eCited by top-tier papers17
- Universal Cross-Tokenizer Distillation via Approximate Likelihood MatchingBenjamin Minixhofer, Ivan Vulic, Edoardo Maria PontiNeurIPS 2025 · 48 citations
- Sampling from Your Language Model One Byte at a TimeJonathan Hayase, Alisa Liu, Noah Smith, Sewoong OhICML 2026 · 9 citations
- The Impact of Token Granularity on the Predictive Power of Language Model SurprisalByung-Doh Oh, William SchulerACL 2025 · 7 citations
- Tokenisation is NP-CompletePhilip Whittington, Gregor Bachmann, Tiago PimentelACL 2025 · 6 citations
- Cross-Tokenizer Likelihood Scoring Algorithms for Language Model DistillationBuu Phan, Ashish Khisti, Karen UllrichICLR 2026 · 5 citations
Builds on5
- Tokenization and the Noiseless ChannelVilém Zouhar, Clara Meister, Juan Luis Gastaldi, Li Du et al.ACL 2023 · 10 citations
- A Measure-Theoretic Characterization of Tight Language ModelsLi Du, Lucas Torroba Hennigen, Tiago Pimentel, Clara Meister et al.ACL 2023 · 8 citations
- You should evaluate your language model on marginal likelihood over tokenisationsKris Cao, Laura RimellEMNLP 2021 · 6 citations
- Revisiting the Optimality of Word LengthsTiago Pimentel, Clara Meister, Ethan Wilcox, Kyle Mahowald et al.EMNLP 2023 · 5 citations
- On the Proper Treatment of Tokenization in PsycholinguisticsMario Giulianelli, Luca Malagutti, Juan Luis Gastaldi, Brian DuSell et al.EMNLP 2024 · 2 citations
Related papers
- Clozing the Gap: Exploring Why Language Model Surprisal Outperforms Cloze SurprisalSathvik Nair, Byung-Doh OhACL 2026 · 1 citation
- Temperature-scaling surprisal estimates improve fit to human reading times - but does it do so for the "right reasons"?Tong Liu, Iza Skrjanec, Vera DembergACL 2024 · 3 citations
- Causal Estimation of Tokenisation BiasPietro Lesci, Clara Meister, Thomas Hofmann, Andreas Vlachos et al.ACL 2025
- Are Language Models Any Good at Density Modeling?Sriram Ranga, Sai Shashank Bedampeta, Rui Mao, Anupam ChattopadhyayAAAI 2026
- Language Model Probabilities are Not Calibrated in Numeric ContextsCharles Lovering, Michael Krumdick, Viet Dac Lai, Varshini Reddy et al.ACL 2025 · 10 citations
