Reasoning Structure of Large Language Models
Frédéric Berdoz, Luca Lanzendörfer, Fabian Farestam, Roger Wattenhofer
Abstract
Large reasoning models (LRMs) are often evaluated using metrics such as final-answer accuracy or token count. However, identical scores on these metrics can hide fundamentally different reasoning structures. To address this limitation, we introduce a scalable LRM benchmark of logic puzzles and a pipeline that converts unstructured traces into verifiable reasoning graphs of claims and dependencies. This turns reasoning into a structured, measurable object whose topology can be quantitatively analyzed. Building on this, we define a reasoning efficiency metric that quantifies how concentrated the model's logical flow is. Our analysis on open-source reasoning models shows that structural measurements separate behaviors that token count and accuracy conflate, providing a practical tool for diagnosing failure modes and comparing how reasoning scales with puzzle difficulty.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bfc8154c-fd70-48a9-8cc9-21e528f13765Builds on16
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- Graph of Thoughts: Solving Elaborate Problems with Large Language ModelsMaciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger et al.AAAI 2024 · 1,292 citations
- The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem ComplexityParshin Shojaee, Iman Mirzadeh, Keivan Alizadeh-Vahid, Maxwell Horton et al.NeurIPS 2025 · 507 citations
- SmartPlay : A Benchmark for LLMs as Intelligent AgentsYue Wu, Xuan Tang, Tom M. Mitchell, Yuanzhi LiICLR 2024 · 121 citations
Related papers
- Mapping the Minds of LLMs: A Graph-Based Analysis of Reasoning LLMsZhen Xiong, Yujun Cai, Zhecheng Li, Yiwei WangEMNLP 2025
- A Temporal Reasoning Benchmarking Framework for LRMs via Difficulty-Controlled and Dynamic Test GenerationShide Zhou, Kailong Wang, Ling Shi, Haoyu WangISSTA 2026
- Characterizing, Evaluating, and Optimizing Complex ReasoningHaoran Zhang, Yafu Li, Zhi Wang, Zhilin Wang et al.ICML 2026 · 2 citations
- DAG-Math: Graph-of-Thought Guided Mathematical Reasoning in LLMsYuanhe Zhang, Ilja Kuzborskij, Jason D. Lee, Chenlei Leng et al.ICLR 2026 · 6 citations
- RPTS: Tree-Structured Reasoning Process Scoring for Faithful Multimodal EvaluationHaofeng Wang, Yu ZhangAAAI 2026
