A Dual-Task Paradigm to Investigate Sentence Comprehension Strategies in Language Models
Rei Emura, Saku Sugawara
Abstract
Language models (LMs) behave more like humans when their cognitive resources are restricted, particularly in predicting sentence processing costs such as reading times. However, it remains unclear whether such constraints similarly affect sentence comprehension strategies. Besides, existing methods do not directly target the balance between memory storage and sentence processing, which is central to human working memory. To address this issue, we propose a dual-task paradigm that combines an arithmetic computation task with a sentence comprehension task, such as "The 2 cocktail + blended 3 =..." Our experiments show that under dual-task conditions, GPT-4o, o3-mini, and o4-mini shift toward plausibility-based comprehension, mirroring humans' rational inference. Specifically, these models show a greater accuracy gap between plausible sentences (e.g., "The cocktail was blended by the bartender") and implausible sentences (e.g., "The bartender was blended by the cocktail") in the dual-task condition compared to the single-task conditions. These findings suggest that constraints on the balance between memory and processing resources promote rational inference in LMs. More broadly, they support the view that human-like sentence comprehension fundamentally arises from the allocation of limited cognitive resources.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9bee8b5f-921a-4fa1-9818-32619ad7ca93Builds on10
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formattingMelanie Sclar, Yejin Choi, Yulia Tsvetkov, Alane SuhrICLR 2024 · 682 citations
- OLMo: Accelerating the Science of Language ModelsDirk Groeneveld, Iz Beltagy, Evan Pete Walsh, Akshita Bhagia et al.ACL 2024 · 52 citations
- Look at the First Sentence: Position Bias in Question AnsweringMiyoung Ko, Jinhyuk Lee, Hyunjae Kim, Gangwoo Kim et al.EMNLP 2020 · 51 citations
- Working Memory Capacity of ChatGPT: An Empirical StudyDongyu Gong, Xingchen Wan, Dingmin WangAAAI 2024 · 31 citations
Related papers
- Memory efficiency and resource-rational encoding in sentence processingWeijie Xu, Brian Dillon, Richard FutrellACL 2026
- Comparing human and language models sentence processing difficulties on complex structuresSamuel Joseph Amouyal, Aya Meltzer-Asscher, Jonathan BerantACL 2026 · 1 citation
- Language Models Trained to do Arithmetic Predict Human Risky and Intertemporal ChoiceJian-Qiao Zhu, Haijiang Yan, Thomas L. GriffithsICLR 2025
- Rollout Roulette: A Probabilistic Inference Approach to Inference-Time Scaling of LLMs using Particle-Based Monte Carlo MethodsIsha Puri, Shivchander Sudalairaj, Guangxuan Xu, Abhishek Bhandwaldar et al.NeurIPS 2025 · 8 citations
- Resource-Rational Noisy-Channel Language Processing: Testing the Effect of Algorithmic Constraints on InferencesThomas Hikaru Clark, Jacob Hoover Vigly, Edward Gibson, Roger P. LevyEMNLP 2025
