You should evaluate your language model on marginal likelihood over tokenisations
Kris Cao, Laura Rimell
摘要
Neural language models typically tokenise input text into sub-word units to achieve an open vocabulary. The standard approach is to use a single canonical tokenisation at both train and test time. We suggest that this approach is unsatisfactory and may bottleneck our evaluation of language model performance. Using only the one-best tokenisation ignores tokeniser uncertainty over alternative tokenisations, which may hurt model out-of-domain performance. In this paper, we argue that instead, language models should be evaluated on their marginal likelihood over tokenisations. We compare different estimators for the marginal likelihood based on sampling, and show that it is feasible to estimate the marginal likelihood with a manageable number of samples. We then evaluate pretrained English and German language models on both the one-besttokenisation and marginal perplexities, and show that the marginal perplexity can be significantly better than the one best, especially on out-of-domain data. We link this difference in perplexity to the tokeniser uncertainty as measured by tokeniser entropy. We discuss some implications of our results for language model training and evaluation, particularly with regard to tokenisation robustness.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Broken Tokens? Your Language Model can Secretly Handle Non-Canonical TokenizationsBrian Siyuan Zheng, Alisa Liu, Orevaoghene Ahia, Jonathan Hayase 等NeurIPS 2025 · 被引用 19 次
- Is Your LLM Overcharging You? Tokenization, Transparency, and IncentivesAnder Artola Velasco, Stratis Tsirtsis, Nastaran Okati, Manuel Gomez-RodriguezICML 2026 · 被引用 16 次
- Sampling from Your Language Model One Byte at a TimeJonathan Hayase, Alisa Liu, Noah Smith, Sewoong OhICML 2026 · 被引用 9 次
- Proxy Compression for Language ModelingLin Zheng, Li Xinyu, Qian Liu, Xiachong Feng 等ICML 2026 · 被引用 3 次
- When to Ensemble: Identifying Token-Level Points for Stable and Fast LLM EnsemblingHeecheol Yun, Kwangmin Ki, Jung Hyun Lee, Eunho YangICLR 2026 · 被引用 3 次
它引用的顶会 Paper3
- Dynamic Programming Encoding for Subword Segmentation in Neural Machine TranslationXuanli He, Gholamreza Haffari, Mohammad NorouziACL 2020 · 被引用 33 次
- BPE-Dropout: Simple and Effective Subword RegularizationIvan Provilkov, Dmitrii Emelianenko, Elena VoitaACL 2020 · 被引用 17 次
- From SPMRL to NMRL: What Did We Learn (and Unlearn) in a Decade of Parsing Morphologically-Rich Languages (MRLs)?Reut Tsarfaty, Dan Bareket, Stav Klein, Amit SekerACL 2020 · 被引用 2 次
相关 Paper
- Where is the signal in tokenization space?Renato Lui Geh, Honghua Zhang, Kareem Ahmed, Benjie Wang 等EMNLP 2024 · 被引用 1 次
- Causal Estimation of Tokenisation BiasPietro Lesci, Clara Meister, Thomas Hofmann, Andreas Vlachos 等ACL 2025
- Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language ModelsPit Neitemeier, Björn Deiseroth, Constantin Eichenberg, Lukas BallesICLR 2025
- How to Compute the Probability of a WordTiago Pimentel, Clara MeisterEMNLP 2024 · 被引用 2 次
- Beyond Text Compression: Evaluating Tokenizers Across ScalesJonas F. Lotz, António Vilarinho Lopes, Stephan Peitz, Hendra Setiawan 等ACL 2025 · 被引用 3 次
