Towards A Scanpath-Conditioned Surprisal Theory: Modeling Reader Information States
Michael Mooney, Edmond S. L. Ho
Abstract
Standard surprisal is typically computed from the linear text prefix, but human reading is non-linear and memory-constrained: readers skip words, regress, and do not retain prior context perfectly. We propose a formulation of surprisal conditioned on a reader-specific accessible information state given by the scanpath history and memory dynamics, rather than by the written prefix alone. Prior context is treated as only probabilistically accessible at each fixation, allowing predictability to depend on both non-linear exposure and forgetting. We evaluate the approach on eye-tracking corpora using held-out log-likelihood over standard duration-based reading measures. Across model variants, conditioning on accessible information states improves predictive fit over standard surprisal baselines. These results suggest that predictability in human reading is better characterized relative to the reader’s evolving accessible information state than to the written prefix alone.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 16b32a51-3690-4ede-acc3-4ff81a0cbc3bBuilds on7
- Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and InferenceBenjamin Warner, Antoine Chaffin, Benjamin Clavié, Orion Weller et al.ACL 2025 · 552 citations
- Masked Language Modeling and the Distributional Hypothesis: Order Word Matters Pre-training for LittleKoustuv Sinha, Robin Jia, Dieuwke Hupkes, Joelle Pineau et al.EMNLP 2021 · 177 citations
- Masked Language Model ScoringJulian Salazar, Davis Liang, Toan Q. Nguyen, Katrin KirchhoffACL 2020 · 167 citations
- The Impact of Token Granularity on the Predictive Power of Language Model SurprisalByung-Doh Oh, William SchulerACL 2025 · 7 citations
- Information Value: Measuring Utterance Predictability as Distance from Plausible AlternativesMario Giulianelli, Sarenne Wallbridge, Raquel FernándezEMNLP 2023 · 4 citations
Related papers
- A Spatio-Temporal Point Process for Fine-Grained Modeling of Reading BehaviorFrancesco Ignazio Re, Andreas Opedal, Glib Manaiev, Mario Giulianelli et al.ACL 2025 · 2 citations
- Probing for Reading TimesEleftheria Tsipidi, Samuel Kiegeland, Francesco Ignazio Re, Tianyang Xu et al.ACL 2026
- Déjà Vu? Decoding Repeated Reading from Eye MovementsYoav Meiri, Omer Shubi, Cfir Avraham Hadar, Ariel Kreisberg Nitzav et al.ACL 2025 · 1 citation
- Entropy- and Distance-Based Predictors From GPT-2 Attention Patterns Predict Reading Times Over and Above GPT-2 SurprisalByung-Doh Oh, William SchulerEMNLP 2022 · 13 citations
- Suspense in Short Stories is Predicted By Uncertainty Reduction over Neural Story RepresentationDavid Wilmot, Frank KellerACL 2020 · 2 citations
