How to Compute the Probability of a Word
Tiago Pimentel, Clara Meister
摘要
Language models (LMs) estimate a probability distribution over strings in a natural language; these distributions are crucial for computing perplexity and surprisal in linguistics research. While we are usually concerned with measuring these values for words, most LMs operate over subwords. Despite seemingly straightforward, accurately computing probabilities over one unit given probabilities over the other requires care. Indeed, we show here that many recent linguistic studies have been incorrectly computing these values. This paper derives the correct methods for computing word probabilities, highlighting issues when relying on language models that use beginning-of-word (bow)-marking tokenisers, e.g., the GPT family. Empirically, we show that correcting the widespread bug in probability computations affects measured outcomes in sentence comprehension and lexical optimisation analyses. tpimentelms/probability-of-a-word pip install wordsprobability
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- Universal Cross-Tokenizer Distillation via Approximate Likelihood MatchingBenjamin Minixhofer, Ivan Vulic, Edoardo Maria PontiNeurIPS 2025 · 被引用 48 次
- Sampling from Your Language Model One Byte at a TimeJonathan Hayase, Alisa Liu, Noah Smith, Sewoong OhICML 2026 · 被引用 9 次
- The Impact of Token Granularity on the Predictive Power of Language Model SurprisalByung-Doh Oh, William SchulerACL 2025 · 被引用 7 次
- Tokenisation is NP-CompletePhilip Whittington, Gregor Bachmann, Tiago PimentelACL 2025 · 被引用 6 次
- Cross-Tokenizer Likelihood Scoring Algorithms for Language Model DistillationBuu Phan, Ashish Khisti, Karen UllrichICLR 2026 · 被引用 5 次
它引用的顶会 Paper5
- Tokenization and the Noiseless ChannelVilém Zouhar, Clara Meister, Juan Luis Gastaldi, Li Du 等ACL 2023 · 被引用 10 次
- A Measure-Theoretic Characterization of Tight Language ModelsLi Du, Lucas Torroba Hennigen, Tiago Pimentel, Clara Meister 等ACL 2023 · 被引用 8 次
- You should evaluate your language model on marginal likelihood over tokenisationsKris Cao, Laura RimellEMNLP 2021 · 被引用 6 次
- Revisiting the Optimality of Word LengthsTiago Pimentel, Clara Meister, Ethan Wilcox, Kyle Mahowald 等EMNLP 2023 · 被引用 5 次
- On the Proper Treatment of Tokenization in PsycholinguisticsMario Giulianelli, Luca Malagutti, Juan Luis Gastaldi, Brian DuSell 等EMNLP 2024 · 被引用 2 次
相关 Paper
- Clozing the Gap: Exploring Why Language Model Surprisal Outperforms Cloze SurprisalSathvik Nair, Byung-Doh OhACL 2026 · 被引用 1 次
- Temperature-scaling surprisal estimates improve fit to human reading times - but does it do so for the "right reasons"?Tong Liu, Iza Skrjanec, Vera DembergACL 2024 · 被引用 3 次
- Causal Estimation of Tokenisation BiasPietro Lesci, Clara Meister, Thomas Hofmann, Andreas Vlachos 等ACL 2025
- Are Language Models Any Good at Density Modeling?Sriram Ranga, Sai Shashank Bedampeta, Rui Mao, Anupam ChattopadhyayAAAI 2026
- Language Model Probabilities are Not Calibrated in Numeric ContextsCharles Lovering, Michael Krumdick, Viet Dac Lai, Varshini Reddy 等ACL 2025 · 被引用 10 次
