A Targeted Assessment of Incremental Processing in Neural Language Models and Humans
Ethan Wilcox, Pranali Vani, Roger Levy
Abstract
We present a targeted, scaled-up comparison of incremental processing in humans and neural language models by collecting by-word reaction time data for sixteen different syntactic test suites across a range of structural phenomena. Human reaction time data comes from a novel online experimental paradigm called the Interpolated Maze task. We compare human reaction times to by-word probabilities for four contemporary language models, with different architectures and trained on a range of data set sizes. We find that across many phenomena, both humans and language models show increased processing difficulty in ungrammatical sentence regions with human and model 'accuracy' scores (à la Marvin and Linzen (2018)) about equal. However, although language model outputs match humans in direction, we show that models systematically under-predict the difference in magnitude of incremental processing difficulty between grammatical and ungrammatical sentences. Specifically, when models encounter syntactic violations they fail to accurately predict the longer reaction times observed in the human data. These results call into question whether contemporary language models are approaching human-like performance for sensitivity to syntactic violations. 1. S d (was impressed) < Sc(was impressed) 2. S d (was impressed) < S b (was impressed) 3. (S d (was impressed) -Sc(was impressed)) < (S b (was impressed) -Sa(was impressed))
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7b62dc78-4ed2-4e28-a93a-95455c5d9216Cited by top-tier papers5
- Neural reality of argument structure constructionsBai Li, Zining Zhu, Guillaume Thomas, Frank Rudzicz et al.ACL 2022 · 38 citations
- Structural Priming Demonstrates Abstract Grammatical Representations in Multilingual Language ModelsJames A. Michaelov, Catherine Arnett, Tyler A. Chang, Ben BergenEMNLP 2023 · 6 citations
- Dual Alignment Between Language Model Layers and Human Sentence ProcessingTatsuki Kuribayashi, Alex Warstadt, Yohei Oseki, Ethan Gotlieb WilcoxACL 2026 · 1 citation
- On the Effect of Hyperparameters in Language Modeling for Computational LinguisticsRuoxi Ning, Yongpeng Zhu, Qingcheng Zeng, Tatsuki Kuribayashi et al.ACL 2026
- When Only Time Will Tell: Interpreting How Transformers Process Local Ambiguities Through the Lens of Restart-IncrementalityBrielen Madureira, Patrick Kahardipraja, David SchlangenACL 2024
Builds on1
Related papers
- Multipath parsing in the brainBerta Franzluebbers, Donald Dunagan, Milos Stanojevic, Jan Buys et al.ACL 2024 · 1 citation
- Clozing the Gap: Exploring Why Language Model Surprisal Outperforms Cloze SurprisalSathvik Nair, Byung-Doh OhACL 2026 · 1 citation
- Comparing human and language models sentence processing difficulties on complex structuresSamuel Joseph Amouyal, Aya Meltzer-Asscher, Jonathan BerantACL 2026 · 1 citation
- When the LM misunderstood the human chuckled: Analyzing garden path effects in humans and language modelsSamuel Joseph Amouyal, Aya Meltzer-Asscher, Jonathan BerantACL 2025
- An Existence Proof for Neural Language Models That Can Explain Garden-Path Effects via SurprisalRyo Yoshida, Shinnosuke Isono, Taiga Someya, Yohei Oseki et al.ACL 2026
