Clozing the Gap: Exploring Why Language Model Surprisal Outperforms Cloze Surprisal
Sathvik Nair, Byung-Doh Oh
Abstract
How predictable a word is can be generally quantified in two ways: using human responses to the cloze task or using probabilities from language models (LMs). When used as predictors of processing effort, LM probabilities outperform probabilities derived from cloze data. However, it is important to establish that LM probabilities do so for the right reasons, since different predictors can lead to different scientific conclusions about the role of prediction in language comprehension. We present evidence for three hypotheses about the apparent advantage of LM probabilities: not suffering from low resolution, distinguishing semantically similar words, and accurately assigning probabilities to low-frequency words. These results call for efforts to improve the resolution of cloze studies, coupled with experiments on whether human-like prediction is also as sensitive to the fine-grained distinctions made by LM probabilities.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3dd532f3-cba5-44cf-9fc6-8b72d61dead3Builds on4
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
- Revisiting the Uniform Information Density HypothesisClara Meister, Tiago Pimentel, Patrick Haller, Lena A. Jäger et al.EMNLP 2021 · 4 citations
- How to Compute the Probability of a WordTiago Pimentel, Clara MeisterEMNLP 2024 · 2 citations
- On the Proper Treatment of Tokenization in PsycholinguisticsMario Giulianelli, Luca Malagutti, Juan Luis Gastaldi, Brian DuSell et al.EMNLP 2024 · 2 citations
Related papers
- Sorting through the noise: Testing robustness of information processing in pre-trained language modelsLalchand Pandia, Allyson EttingerEMNLP 2021 · 20 citations
- A Targeted Assessment of Incremental Processing in Neural Language Models and HumansEthan Wilcox, Pranali Vani, Roger LevyACL 2021
- Temperature-scaling surprisal estimates improve fit to human reading times - but does it do so for the "right reasons"?Tong Liu, Iza Skrjanec, Vera DembergACL 2024 · 3 citations
- Dual Alignment Between Language Model Layers and Human Sentence ProcessingTatsuki Kuribayashi, Alex Warstadt, Yohei Oseki, Ethan Gotlieb WilcoxACL 2026 · 1 citation
- Rolling the DICE on Idiomaticity: How LLMs Fail to Grasp ContextMaggie Mi, Aline Villavicencio, Nafise Sadat MoosaviACL 2025 · 9 citations
