Step-by-Step Reasoning to Solve Grid Puzzles: Where do LLMs Falter?
Nemika Tyagi, Mihir Parmar, Mohith Kulkarni, Aswin RRV, Nisarg Patel, Mutsumi Nakamura, Arindam Mitra, Chitta Baral
摘要
Solving grid puzzles involves a significant amount of logical reasoning. Hence, it is a good domain to evaluate the reasoning capability of a model which can then guide us to improve the reasoning ability of models. However, most existing works evaluate only the final predicted answer of a puzzle, without delving into an indepth analysis of the LLMs' reasoning chains (such as where they falter) or providing any finer metrics to evaluate them. Since LLMs may rely on simple heuristics or artifacts to predict the final answer, it is crucial to evaluate the generated reasoning chain beyond overall correctness measures, for accurately evaluating the reasoning abilities of LLMs. To this end, we first develop GridPuzzle, an evaluation dataset comprising 274 grid-based puzzles with different complexities. Second, we propose a new error taxonomy derived from manual analysis of reasoning chains from LLMs including GPT-4, Claude-3, Gemini, Mistral, and Llama-2. Then, we develop an LLM-based framework for large-scale subjective evaluation (i.e., identifying errors) and an objective metric, PuzzleEval, to evaluate the correctness of reasoning chains. Evaluating reasoning chains from LLMs leads to several interesting findings. We further show that existing prompting methods used for enhancing models' reasoning abilities do not improve performance on Grid-Puzzle. This highlights the importance of understanding fine-grained errors and presents a challenge for future research to enhance LLMs' puzzle-solving abilities by developing methods that address these errors 1 . x 5 4 x 4 4 x 5 4 x 6 x 4 A group of friends has decided to try several different weight-loss diets and exercises to see who amongst them can lose the most weight in 3 months. Using only the clues below, match the pounds lost to the options from names and diets. Remember, as with all grid-based logic puzzles, no option in any category will ever be used more than once.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Probabilistic Soundness Guarantees in LLM Reasoning ChainsWeiqiu You, Anton Xue, Shreya Havaldar, Delip Rao 等EMNLP 2025 · 被引用 9 次
- A Solver-in-the-Loop Framework for Improving LLMs on Answer Set Programming for Logic Puzzle SolvingTimo Pierre Schrader, Lukas Lange, Tobias Kaminski, Simon Razniewski 等AAAI 2026 · 被引用 4 次
- FineReason: Evaluating and Improving LLMs' Deliberate Reasoning through Reflective Puzzle SolvingGuizhen Chen, Weiwen Xu, Hao Zhang, Hou Pong Chan 等ACL 2025
- Lexical Recall or Logical Reasoning: Probing the Limits of Reasoning Abilities in Large Language ModelsHenrike Beyer, Chris ReedACL 2025
- ZebraLogic: On the Scaling Limits of LLMs for Logical ReasoningBill Yuchen Lin, Ronan Le Bras, Kyle Richardson, Ashish Sabharwal 等ICML 2025
它引用的顶会 Paper15
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
- BARTScore: Evaluating Generated Text as Text GenerationWeizhe Yuan, Graham Neubig, Pengfei LiuNeurIPS 2021 · 被引用 1,143 次
- G-Eval: NLG Evaluation using Gpt-4 with Better Human AlignmentYang Liu, Dan Iter, Yichong Xu, Shuohang Wang 等EMNLP 2023 · 被引用 549 次
- Can Large Language Models Be an Alternative to Human Evaluations?David Cheng-Han Chiang, Hung-yi LeeACL 2023 · 被引用 254 次
相关 Paper
- LogicBench: Towards Systematic Evaluation of Logical Reasoning Ability of Large Language ModelsMihir Parmar, Nisarg Patel, Neeraj Varshney, Mutsumi Nakamura 等ACL 2024
- Conditional and Modal Reasoning in Large Language ModelsWesley H. Holliday, Matthew Mandelkern, Cedegao ZhangEMNLP 2024 · 被引用 5 次
- Puzzle Solving using Reasoning of Large Language Models: A SurveyPanagiotis Giadikiaroglou, Maria Lymperaiou, Giorgos Filandrianos, Giorgos StamouEMNLP 2024 · 被引用 9 次
- Exposing the Achilles' Heel: Evaluating LLMs Ability to Handle Mistakes in Mathematical ReasoningJoykirat Singh, Akshay Uttama Nambi, Vibhav VineetACL 2025 · 被引用 10 次
- No Need for Explanations: LLMs can implicitly learn from mistakes in-contextLisa Alazraki, Maximilian Mozes, Jon Ander Campos, Yi Chern Tan 等EMNLP 2025
