A Dual-Task Paradigm to Investigate Sentence Comprehension Strategies in Language Models
Rei Emura, Saku Sugawara
摘要
Language models (LMs) behave more like humans when their cognitive resources are restricted, particularly in predicting sentence processing costs such as reading times. However, it remains unclear whether such constraints similarly affect sentence comprehension strategies. Besides, existing methods do not directly target the balance between memory storage and sentence processing, which is central to human working memory. To address this issue, we propose a dual-task paradigm that combines an arithmetic computation task with a sentence comprehension task, such as "The 2 cocktail + blended 3 =..." Our experiments show that under dual-task conditions, GPT-4o, o3-mini, and o4-mini shift toward plausibility-based comprehension, mirroring humans' rational inference. Specifically, these models show a greater accuracy gap between plausible sentences (e.g., "The cocktail was blended by the bartender") and implausible sentences (e.g., "The bartender was blended by the cocktail") in the dual-task condition compared to the single-task conditions. These findings suggest that constraints on the balance between memory and processing resources promote rational inference in LMs. More broadly, they support the view that human-like sentence comprehension fundamentally arises from the allocation of limited cognitive resources.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formattingMelanie Sclar, Yejin Choi, Yulia Tsvetkov, Alane SuhrICLR 2024 · 被引用 682 次
- OLMo: Accelerating the Science of Language ModelsDirk Groeneveld, Iz Beltagy, Evan Pete Walsh, Akshita Bhagia 等ACL 2024 · 被引用 52 次
- Look at the First Sentence: Position Bias in Question AnsweringMiyoung Ko, Jinhyuk Lee, Hyunjae Kim, Gangwoo Kim 等EMNLP 2020 · 被引用 51 次
- Working Memory Capacity of ChatGPT: An Empirical StudyDongyu Gong, Xingchen Wan, Dingmin WangAAAI 2024 · 被引用 31 次
相关 Paper
- Memory efficiency and resource-rational encoding in sentence processingWeijie Xu, Brian Dillon, Richard FutrellACL 2026
- Comparing human and language models sentence processing difficulties on complex structuresSamuel Joseph Amouyal, Aya Meltzer-Asscher, Jonathan BerantACL 2026 · 被引用 1 次
- Language Models Trained to do Arithmetic Predict Human Risky and Intertemporal ChoiceJian-Qiao Zhu, Haijiang Yan, Thomas L. GriffithsICLR 2025
- Rollout Roulette: A Probabilistic Inference Approach to Inference-Time Scaling of LLMs using Particle-Based Monte Carlo MethodsIsha Puri, Shivchander Sudalairaj, Guangxuan Xu, Abhishek Bhandwaldar 等NeurIPS 2025 · 被引用 8 次
- Resource-Rational Noisy-Channel Language Processing: Testing the Effect of Algorithmic Constraints on InferencesThomas Hikaru Clark, Jacob Hoover Vigly, Edward Gibson, Roger P. LevyEMNLP 2025
