Modeling Hierarchical Thinking in Large Reasoning Models
G M Shahariar, Erfan Shayegani, Ali Nazari, Nael Abu-Ghazaleh
Abstract
Large Reasoning Models (LRMs) solve complex tasks by generating long Chain-of-Thought (CoT) sequences; however, the emergent dynamics governing reasoning trajectories are not well understood and can lead to inconsistencies and reasoning pathologies. In this work, we propose to approximate LRM's emerging hierarchical reasoning dynamics as a trajectory within a Finite State Machine (FSM) transitioning among six abstract cognitive states. We demonstrate that these states and transitions can be captured in the latent state of the model. We believe that this representation can have different applications in the interpretability and optimization of LRM models. For example, by analyzing the topology of these transitions, we identify statistical shifts in reasoning strategies that help identify effective reasoning chains from those that fail. To illustrate these potential advantages, we propose Q-Value guided steering, a training-free inference-time control method that treats reasoning as a planning problem. We estimate the long-horizon utility of state transitions and apply sparse, orthogonal activation steering at sentence boundaries to align the CoT generation with optimal reasoning policies. Experiments across four benchmarks (AIME25, MATH-500, GSM8k, and GPQA Diamond) using three stateof-the-art open reasoning models demonstrate that Q-Value steering policy achieves significant performance gains with "surgical" efficiency, often requiring 25× fewer interventions than greedy and weighted baselines, which suggests that reasoning can be effectively controlled by guiding high-level cognitive dynamics rather than micromanaging token generation. Code is available at: https://github.com/shahariar-shibli/CoT-FSM.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 59bea081-8fae-4256-8848-868cecfcd7bfBuilds on8
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Graph of Thoughts: Solving Elaborate Problems with Large Language ModelsMaciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger et al.AAAI 2024 · 1,292 citations
- Refusal in Language Models Is Mediated by a Single DirectionAndy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka et al.NeurIPS 2024 · 1,166 citations
- RL Squeezes, SFT Expands: A Comparative Study of Reasoning LLMsKohsei Matsutani, Shota Takashiro, Gouki Minegishi, Takeshi Kojima et al.ICLR 2026 · 36 citations
- Topology of Reasoning: Understanding Large Reasoning Models through Reasoning Graph PropertiesGouki Minegishi, Hiroki Furuta, Takeshi Kojima, Yusuke Iwasawa et al.NeurIPS 2025 · 31 citations
Related papers
- Eliciting Chain-of-Thought in Base LLMs via Gradient-Based Representation OptimizationZijian Wang, Yanxiang Ma, Chang XuAAAI 2026
- LLM Reasoning as Trajectories: Step-Specific Representation Geometry and Correctness SignalsLihao Sun, Hang Dong, Bo Qiao, Qingwei Lin et al.ACL 2026 · 9 citations
- DenseSteer: Steering Small Language Models towards Dense Math ReasoningYang Ouyang, Shuhang Lin, Jung-Eun KimICML 2026 · 1 citation
- Unlocking the Black Box of Latent Reasoning: An Interpretability-Guided Approach to InterventionShuochen Chang, Tong Bai, Xiaofeng Zhang, Qianli Ma et al.ACL 2026 · 1 citation
- Feature Extraction and Steering for Enhanced Chain-of-Thought Reasoning in Language ModelsZihao Li, Xu Wang, Yuzhe Yang, Ziyu Yao et al.EMNLP 2025 · 14 citations
