A Targeted Assessment of Incremental Processing in Neural Language Models and Humans
Ethan Wilcox, Pranali Vani, Roger Levy
摘要
We present a targeted, scaled-up comparison of incremental processing in humans and neural language models by collecting by-word reaction time data for sixteen different syntactic test suites across a range of structural phenomena. Human reaction time data comes from a novel online experimental paradigm called the Interpolated Maze task. We compare human reaction times to by-word probabilities for four contemporary language models, with different architectures and trained on a range of data set sizes. We find that across many phenomena, both humans and language models show increased processing difficulty in ungrammatical sentence regions with human and model 'accuracy' scores (à la Marvin and Linzen (2018)) about equal. However, although language model outputs match humans in direction, we show that models systematically under-predict the difference in magnitude of incremental processing difficulty between grammatical and ungrammatical sentences. Specifically, when models encounter syntactic violations they fail to accurately predict the longer reaction times observed in the human data. These results call into question whether contemporary language models are approaching human-like performance for sensitivity to syntactic violations. 1. S d (was impressed) < Sc(was impressed) 2. S d (was impressed) < S b (was impressed) 3. (S d (was impressed) -Sc(was impressed)) < (S b (was impressed) -Sa(was impressed))
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Neural reality of argument structure constructionsBai Li, Zining Zhu, Guillaume Thomas, Frank Rudzicz 等ACL 2022 · 被引用 38 次
- Structural Priming Demonstrates Abstract Grammatical Representations in Multilingual Language ModelsJames A. Michaelov, Catherine Arnett, Tyler A. Chang, Ben BergenEMNLP 2023 · 被引用 6 次
- Dual Alignment Between Language Model Layers and Human Sentence ProcessingTatsuki Kuribayashi, Alex Warstadt, Yohei Oseki, Ethan Gotlieb WilcoxACL 2026 · 被引用 1 次
- On the Effect of Hyperparameters in Language Modeling for Computational LinguisticsRuoxi Ning, Yongpeng Zhu, Qingcheng Zeng, Tatsuki Kuribayashi 等ACL 2026
- When Only Time Will Tell: Interpreting How Transformers Process Local Ambiguities Through the Lens of Restart-IncrementalityBrielen Madureira, Patrick Kahardipraja, David SchlangenACL 2024
它引用的顶会 Paper1
相关 Paper
- Multipath parsing in the brainBerta Franzluebbers, Donald Dunagan, Milos Stanojevic, Jan Buys 等ACL 2024 · 被引用 1 次
- Clozing the Gap: Exploring Why Language Model Surprisal Outperforms Cloze SurprisalSathvik Nair, Byung-Doh OhACL 2026 · 被引用 1 次
- Comparing human and language models sentence processing difficulties on complex structuresSamuel Joseph Amouyal, Aya Meltzer-Asscher, Jonathan BerantACL 2026 · 被引用 1 次
- When the LM misunderstood the human chuckled: Analyzing garden path effects in humans and language modelsSamuel Joseph Amouyal, Aya Meltzer-Asscher, Jonathan BerantACL 2025
- An Existence Proof for Neural Language Models That Can Explain Garden-Path Effects via SurprisalRyo Yoshida, Shinnosuke Isono, Taiga Someya, Yohei Oseki 等ACL 2026
