Working Memory Identifies Reasoning Limits in Language Models
Chunhui Zhang, Yiren Jian, Zhongyu Ouyang, Soroush Vosoughi
摘要
This study explores the inherent limitations of Large Language Models (LLMs) from a scaling perspective, focusing on the upper bounds of their cognitive capabilities. We integrate insights from cognitive science to quantitatively examine how LLMs perform on n-back tasks-a benchmark used to assess working memory, which involves temporarily holding and manipulating information. Our findings reveal that despite increased model size, LLMs still face significant challenges in holding and processing information effectively, especially under complex task conditions. We also assess various prompting strategies, revealing their diverse impacts on LLM performance. The results highlight the struggle of current LLMs to autonomously discover optimal problemsolving patterns without heavily relying on manually corrected prompts. To move beyond these constraints, fundamental improvements in the planning and search of LLMs are essential for them to reason autonomously. Improving these capabilities will reduce the reliance on external corrections and enable LLMs to become more autonomous in their problemsolving processes.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Democratizing Large Language Models via Personalized Parameter-Efficient Fine-tuningZhaoxuan Tan, Qingkai Zeng, Yijun Tian, Zheyuan Liu 等EMNLP 2024 · 被引用 17 次
- Superficial Self-Improved Reasoners Benefit from Model MergingXiangchi Yuan, Chunhui Zhang, Zheyuan Liu, Dachuan Shi 等EMNLP 2025 · 被引用 15 次
- Personalized Pieces: Efficient Personalized Large Language Models through Collaborative EffortsZhaoxuan Tan, Zheyuan Liu, Meng JiangEMNLP 2024 · 被引用 11 次
- On the Eligibility of LLMs for Counterfactual Reasoning: A Decompositional StudyShuai Yang, Qi Yang, Luoxi Tang, Yuqiao Meng 等ICLR 2026 · 被引用 9 次
- Growing Through Experience: Scaling Episodic Grounding in Language ModelsChunhui Zhang, Sirui Wang, Zhongyu Ouyang, Xiangchi Yuan 等ACL 2025 · 被引用 6 次
它引用的顶会 Paper20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
相关 Paper
- Working Memory Capacity of ChatGPT: An Empirical StudyDongyu Gong, Xingchen Wan, Dingmin WangAAAI 2024 · 被引用 31 次
- The Illusion of Diminishing Returns: Measuring Long Horizon Execution in LLMsAkshit Sinha, Arvindh Arun, Shashwat Goel, Steffen Staab 等ICLR 2026 · 被引用 64 次
- ExpeTrans: LLMs Are Experiential Transfer LearnersJinglong Gao, Xiao Ding, Lingxiao Zou, Bibo Cai 等ACL 2025
- LLMs Can Plan Only If We Tell ThemBilgehan Sel, Ruoxi Jia, Ming JinICLR 2025
- SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language ModelsXiaoxuan Wang, Ziniu Hu, Pan Lu, Yanqiao Zhu 等ICML 2024 · 被引用 220 次
