The Impact of Token Granularity on the Predictive Power of Language Model Surprisal
Byung-Doh Oh, William Schuler
摘要
Word-by-word language model surprisal is often used to model the incremental processing of human readers, which raises questions about how various choices in language modeling influence its predictive power. One factor that has been overlooked in cognitive modeling is the granularity of subword tokens, which explicitly encodes information about word length and frequency, and ultimately influences the quality of vector representations that are learned. This paper presents experiments that manipulate the token granularity and evaluate its impact on the ability of surprisal to account for processing difficulty of naturalistic text and garden-path constructions. Experiments with naturalistic reading times reveal a substantial influence of token granularity on surprisal, with tokens defined by a vocabulary size of 8,000 resulting in surprisal that is most predictive. In contrast, on garden-path constructions, language models trained on coarser-grained tokens generally assigned higher surprisal to critical regions, suggesting a greater sensitivity to garden-path effects than previously reported. Taken together, these results suggest a large role of token granularity on the quality of language model surprisal for cognitive modeling.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Dual Alignment Between Language Model Layers and Human Sentence ProcessingTatsuki Kuribayashi, Alex Warstadt, Yohei Oseki, Ethan Gotlieb WilcoxACL 2026 · 被引用 1 次
- Towards A Scanpath-Conditioned Surprisal Theory: Modeling Reader Information StatesMichael Mooney, Edmond S. L. HoACL 2026
- On the Proper Treatment of Units in Surprisal TheorySamuel Kiegeland, Vésteinn Snæbjarnarson, Tim Vieira, Ryan CotterellACL 2026
它引用的顶会 Paper9
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley 等ICML 2023 · 被引用 1,822 次
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space DualityTri Dao, Albert GuICML 2024 · 被引用 1,407 次
- Tokenization Is More Than CompressionCraig W. Schmidt, Varshini Reddy, Haoran Zhang, Alec Alameddine 等EMNLP 2024 · 被引用 16 次
- Tokenization and the Noiseless ChannelVilém Zouhar, Clara Meister, Juan Luis Gastaldi, Li Du 等ACL 2023 · 被引用 10 次
- Fishing for Magikarp: Automatically Detecting Under-trained Tokens in Large Language ModelsSander Land, Max BartoloEMNLP 2024 · 被引用 4 次
相关 Paper
- An Existence Proof for Neural Language Models That Can Explain Garden-Path Effects via SurprisalRyo Yoshida, Shinnosuke Isono, Taiga Someya, Yohei Oseki 等ACL 2026
- Temperature-scaling surprisal estimates improve fit to human reading times - but does it do so for the "right reasons"?Tong Liu, Iza Skrjanec, Vera DembergACL 2024 · 被引用 3 次
- On the Proper Treatment of Tokenization in PsycholinguisticsMario Giulianelli, Luca Malagutti, Juan Luis Gastaldi, Brian DuSell 等EMNLP 2024 · 被引用 2 次
- How to Compute the Probability of a WordTiago Pimentel, Clara MeisterEMNLP 2024 · 被引用 2 次
- When the LM misunderstood the human chuckled: Analyzing garden path effects in humans and language modelsSamuel Joseph Amouyal, Aya Meltzer-Asscher, Jonathan BerantACL 2025
