FineReason: Evaluating and Improving LLMs' Deliberate Reasoning through Reflective Puzzle Solving
Guizhen Chen, Weiwen Xu, Hao Zhang, Hou Pong Chan, Chaoqun Liu, Lidong Bing, Deli Zhao, Anh Tuan Luu, Yu Rong
Abstract
Many challenging reasoning tasks require not just rapid, intuitive responses, but a more deliberate, multi-step approach. Recent progress in large language models (LLMs) highlights an important shift from the "System 1" way of quick reactions to the "System 2" style of reflection-and-correction problem solving. However, current benchmarks heavily rely on the final-answer accuracy, leaving much of a model's intermediate reasoning steps unexamined. This fails to assess the model's ability to reflect and rectify mistakes within the reasoning process. To bridge this gap, we introduce FINEREASON, a logic-puzzle benchmark for systematic evaluation of LLMs' reasoning capabilities. Each puzzle can be decomposed into atomic steps, making it ideal for rigorous validation of intermediate correctness. Building on this, we introduce two tasks: state checking and state transition, for a comprehensive evaluation of how models assess the current situation and plan the next move. To support broader research, we also provide a puzzle training set aimed at enhancing general reasoning. We show that models trained on our state checking and transition data demonstrate gains in mathematical reasoning by up to 5.1%. 1 * Guizhen and Chaoqun are under the Joint PhD Program between Alibaba and NTU. † Corresponding authors. 1 https://github.com/DAMO-NLP-SG/FineReason … (some steps skipped)
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- PuzzleWorld: A Benchmark for Multimodal, Open-Ended Reasoning in PuzzlehuntsHengzhi Li, Justin Zhang, Brendon Jiang, Alexander Naehu et al.ICLR 2026 · 6 citations
- ReasonMed: A 370K Multi-Agent Generated Dataset for Advancing Medical ReasoningYu Sun, Xingyu Qian, Weiwen Xu, Hao Zhang et al.EMNLP 2025 · 1 citation
- SATBench: Benchmarking LLMs' Logical Reasoning via Automated Puzzle Generation from SAT FormulasAnjiang Wei, Yuheng Wu, Yingjia Wan, Tarun Suresh et al.EMNLP 2025 · 1 citation
Builds on12
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- STaR: Bootstrapping Reasoning With ReasoningEric Zelikman, Yuhuai Wu, Jesse Mu, Noah D. GoodmanNeurIPS 2022 · 1,126 citations
- Faith and Fate: Limits of Transformers on CompositionalityNouha Dziri, Ximing Lu, Melanie Sclar, Xiang Lorraine Li et al.NeurIPS 2023 · 728 citations
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le et al.ICLR 2023 · 681 citations
Related papers
- MR-Ben: A Meta-Reasoning Benchmark for Evaluating System-2 Thinking in LLMsZhongshen Zeng, Yinhong Liu, Yingjia Wan, Jingyao Li et al.NeurIPS 2024 · 51 citations
- LogicBench: Towards Systematic Evaluation of Logical Reasoning Ability of Large Language ModelsMihir Parmar, Nisarg Patel, Neeraj Varshney, Mutsumi Nakamura et al.ACL 2024
- ALERT: Adapt Language Models to Reasoning TasksPing Yu, Tianlu Wang, Olga Golovneva, Badr AlKhamissi et al.ACL 2023 · 4 citations
- VisualPuzzles: Decoupling Multimodal Reasoning Evaluation from Domain KnowledgeYueqi Song, Tianyue Ou, Yibo Kong, Zecheng Li et al.ICML 2026 · 44 citations
- RefineBench: Evaluating Refinement Capability of Language Models via ChecklistsYoung-Jun Lee, Seungone Kim, Byung-Kwan Lee, Minkyeong Moon et al.ICLR 2026 · 13 citations
