Lexical Recall or Logical Reasoning: Probing the Limits of Reasoning Abilities in Large Language Models
Henrike Beyer, Chris Reed
Abstract
Despite the increasing interest in the reasoning abilities of Large Language Models (LLMs), existing work shows limitations in assessing logic abilities independently from lexical memory. We address this gap with Mystery-Zebra. This robust two-part benchmark (4,290 puzzles) challenges the logic abstraction abilities of LLMs in two setups: (1) a lexical obfuscation setup tests the dependence of LLMs on lexical content based on two canonical grid puzzles widely spread on the Internet; (2) a set of new grid puzzles in 42 different sizes and 12 difficulty levels tests how the formal difficulty degree of a puzzle affects LLMs. We test open and closed-weight LLMs on both parts of the benchmark. The results on part two suggest that model sizes up to 70B parameters have only a minor influence when solving newly generated puzzles, while performance mainly relates to the number of items in the puzzle. The results on the first part of the benchmark suggest that the applied obfuscation strategies help to mitigate effects of logic puzzles being part of LLM training data, showing a drastic drop in performance for obfuscated versions of well-known puzzles. In addition we conduct a case-study on the first part of the benchmark predicting the position of single items, unveiling that the reasoning abilities of LLMs are mainly limited to a few consecutive steps of reasoning. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 57044e3f-ae12-40eb-8c70-17bc18708c3aBuilds on7
- On the Planning Abilities of Large Language Models - A Critical InvestigationKarthik Valmeekam, Matthew Marquez, Sarath Sreedharan, Subbarao KambhampatiNeurIPS 2023 · 509 citations
- Making Them Ask and Answer: Jailbreaking Large Language Models in Few Queries via Disguise and ReconstructionTong Liu, Yingjie Zhang, Zhe Zhao, Yinpeng Dong et al.USENIX Security 2024 · 121 citations
- Causal language modeling can elicit search and reasoning capabilities on logic puzzlesKulin Shah, Nishanth Dikkala, Xin Wang, Rina PanigrahyNeurIPS 2024 · 44 citations
- LogicAsker: Evaluating and Improving the Logical Reasoning Ability of Large Language ModelsYuxuan Wan, Wenxuan Wang, Yiliu Yang, Youliang Yuan et al.EMNLP 2024 · 10 citations
- Puzzle Solving using Reasoning of Large Language Models: A SurveyPanagiotis Giadikiaroglou, Maria Lymperaiou, Giorgos Filandrianos, Giorgos StamouEMNLP 2024 · 9 citations
Related papers
- ZebraLogic: On the Scaling Limits of LLMs for Logical ReasoningBill Yuchen Lin, Ronan Le Bras, Kyle Richardson, Ashish Sabharwal et al.ICML 2025
- Logic.py: Bridging the Gap between LLMs and Constraint SolversPascal Kesseli, Peter W. O'Hearn, Ricardo Silveira CabralNeurIPS 2025 · 10 citations
- SATBench: Benchmarking LLMs' Logical Reasoning via Automated Puzzle Generation from SAT FormulasAnjiang Wei, Yuheng Wu, Yingjia Wan, Tarun Suresh et al.EMNLP 2025 · 1 citation
- Benchmarking Abstract and Reasoning Abilities Through A Theoretical PerspectiveQingchuan Ma, Yuhang Wu, Xiawu Zheng, Rongrong JiICML 2025
- LogicBench: Towards Systematic Evaluation of Logical Reasoning Ability of Large Language ModelsMihir Parmar, Nisarg Patel, Neeraj Varshney, Mutsumi Nakamura et al.ACL 2024
