Lune

EMNLP2024Top-tier venue

Working Memory Identifies Reasoning Limits in Language Models

Chunhui Zhang, Yiren Jian, Zhongyu Ouyang, Soroush Vosoughi

2024Year
4Citations
11Top-tier citations

Abstract

This study explores the inherent limitations of Large Language Models (LLMs) from a scaling perspective, focusing on the upper bounds of their cognitive capabilities. We integrate insights from cognitive science to quantitatively examine how LLMs perform on n-back tasks-a benchmark used to assess working memory, which involves temporarily holding and manipulating information. Our findings reveal that despite increased model size, LLMs still face significant challenges in holding and processing information effectively, especially under complex task conditions. We also assess various prompting strategies, revealing their diverse impacts on LLM performance. The results highlight the struggle of current LLMs to autonomously discover optimal problemsolving patterns without heavily relying on manually corrected prompts. To move beyond these constraints, fundamental improvements in the planning and search of LLMs are essential for them to reason autonomously. Improving these capabilities will reduce the reliance on external corrections and enable LLMs to become more autonomous in their problemsolving processes.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 7cb87f4f-194d-499c-a0dc-5809fb016e75

Cited by top-tier papers11

Ask how each one uses it

Builds on20

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines