ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning
Bill Yuchen Lin, Ronan Le Bras, Kyle Richardson, Ashish Sabharwal, Radha Poovendran, Peter Clark, Yejin Choi
Abstract
We investigate the logical reasoning capabilities of large language models (LLMs) and their scalability in complex non-monotonic reasoning. To this end, we introduce ZebraLogic, a comprehensive evaluation framework for assessing LLM reasoning performance on logic grid puzzles derived from constraint satisfaction problems (CSPs). Ze-braLogic enables the generation of puzzles with controllable and quantifiable complexity, facilitating a systematic study of the scaling limits of models such as Llama, o1 models, and DeepSeek-R1. By encompassing a broad range of search space complexities and diverse logical constraints, ZebraLogic provides a structured environment to evaluate reasoning under increasing difficulty. Our results reveal a significant decline in accuracy as problem complexity grows-a phenomenon we term the "curse of complexity." This limitation persists even with larger models and increased inference-time computation, suggesting inherent constraints in current LLM reasoning capabilities. Additionally, we explore strategies to enhance logical reasoning, including Best-of-N sampling, backtracking mechanisms, and self-verification prompts. Our findings offer critical insights into the scalability of LLM reasoning, highlight fundamental limitations, and outline potential directions for improvement.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c828d228-1e91-44cb-9af3-d6869fbe5bb0Cited by top-tier papers12
- SLR: Automated Synthesis for Scalable Logical ReasoningLukas Helff, Ahmad Omar, Felix Friedrich, Antonia Wüst et al.ACL 2026 · 6 citations
- Language Model as Planner and Formalizer under ConstraintsCassie Huang, Stuti Mohan, Ziyi Yang, Stefanie Tellex et al.ACL 2026 · 3 citations
- SATBench: Benchmarking LLMs' Logical Reasoning via Automated Puzzle Generation from SAT FormulasAnjiang Wei, Yuheng Wu, Yingjia Wan, Tarun Suresh et al.EMNLP 2025 · 1 citation
- ReEfBench: Quantifying the Reasoning Efficiency of LLMsZhizhang Fu, Yuancheng Gu, Chenkai Hu, Hanmeng Liu et al.ACL 2026 · 1 citation
- Large Language Model for OWL ProofsHui Yang, Jiaoyan Chen, Uli SattlerWWW 2026 · 1 citation
Builds on10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Faith and Fate: Limits of Transformers on CompositionalityNouha Dziri, Ximing Lu, Melanie Sclar, Xiang Lorraine Li et al.NeurIPS 2023 · 728 citations
- Pushing the Limits of Rule Reasoning in Transformers through Natural Language SatisfiabilityKyle Richardson, Ashish SabharwalAAAI 2022 · 29 citations
- Can Transformers Reason in Fragments of Natural Language?Viktor Schlegel, Kamen V. Pavlov, Ian Pratt-HartmannEMNLP 2022 · 4 citations
Related papers
- Lexical Recall or Logical Reasoning: Probing the Limits of Reasoning Abilities in Large Language ModelsHenrike Beyer, Chris ReedACL 2025
- Logic.py: Bridging the Gap between LLMs and Constraint SolversPascal Kesseli, Peter W. O'Hearn, Ricardo Silveira CabralNeurIPS 2025 · 10 citations
- seqBench: A Tunable Benchmark to Quantify Sequential Reasoning Limits of LLMsMohammad Ramezanali, Mo Vazifeh, Paolo SantiEMNLP 2025
- Multi-LogiEval: Towards Evaluating Multi-Step Logical Reasoning Ability of Large Language ModelsNisarg Patel, Mohith Kulkarni, Mihir Parmar, Aashna Budhiraja et al.EMNLP 2024 · 6 citations
- LogicAsker: Evaluating and Improving the Logical Reasoning Ability of Large Language ModelsYuxuan Wan, Wenxuan Wang, Yiliu Yang, Youliang Yuan et al.EMNLP 2024 · 10 citations
