Working Memory Identifies Reasoning Limits in Language Models
Chunhui Zhang, Yiren Jian, Zhongyu Ouyang, Soroush Vosoughi
Abstract
This study explores the inherent limitations of Large Language Models (LLMs) from a scaling perspective, focusing on the upper bounds of their cognitive capabilities. We integrate insights from cognitive science to quantitatively examine how LLMs perform on n-back tasks-a benchmark used to assess working memory, which involves temporarily holding and manipulating information. Our findings reveal that despite increased model size, LLMs still face significant challenges in holding and processing information effectively, especially under complex task conditions. We also assess various prompting strategies, revealing their diverse impacts on LLM performance. The results highlight the struggle of current LLMs to autonomously discover optimal problemsolving patterns without heavily relying on manually corrected prompts. To move beyond these constraints, fundamental improvements in the planning and search of LLMs are essential for them to reason autonomously. Improving these capabilities will reduce the reliance on external corrections and enable LLMs to become more autonomous in their problemsolving processes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7cb87f4f-194d-499c-a0dc-5809fb016e75Cited by top-tier papers11
- Democratizing Large Language Models via Personalized Parameter-Efficient Fine-tuningZhaoxuan Tan, Qingkai Zeng, Yijun Tian, Zheyuan Liu et al.EMNLP 2024 · 17 citations
- Superficial Self-Improved Reasoners Benefit from Model MergingXiangchi Yuan, Chunhui Zhang, Zheyuan Liu, Dachuan Shi et al.EMNLP 2025 · 15 citations
- Personalized Pieces: Efficient Personalized Large Language Models through Collaborative EffortsZhaoxuan Tan, Zheyuan Liu, Meng JiangEMNLP 2024 · 11 citations
- On the Eligibility of LLMs for Counterfactual Reasoning: A Decompositional StudyShuai Yang, Qi Yang, Luoxi Tang, Yuqiao Meng et al.ICLR 2026 · 9 citations
- Growing Through Experience: Scaling Episodic Grounding in Language ModelsChunhui Zhang, Sirui Wang, Zhongyu Ouyang, Xiangchi Yuan et al.ACL 2025 · 6 citations
Builds on20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
Related papers
- Working Memory Capacity of ChatGPT: An Empirical StudyDongyu Gong, Xingchen Wan, Dingmin WangAAAI 2024 · 31 citations
- The Illusion of Diminishing Returns: Measuring Long Horizon Execution in LLMsAkshit Sinha, Arvindh Arun, Shashwat Goel, Steffen Staab et al.ICLR 2026 · 64 citations
- ExpeTrans: LLMs Are Experiential Transfer LearnersJinglong Gao, Xiao Ding, Lingxiao Zou, Bibo Cai et al.ACL 2025
- LLMs Can Plan Only If We Tell ThemBilgehan Sel, Ruoxi Jia, Ming JinICLR 2025
- SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language ModelsXiaoxuan Wang, Ziniu Hu, Pan Lu, Yanqiao Zhu et al.ICML 2024 · 220 citations
