Modeling Hierarchical Thinking in Large Reasoning Models
G M Shahariar, Erfan Shayegani, Ali Nazari, Nael Abu-Ghazaleh
摘要
Large Reasoning Models (LRMs) solve complex tasks by generating long Chain-of-Thought (CoT) sequences; however, the emergent dynamics governing reasoning trajectories are not well understood and can lead to inconsistencies and reasoning pathologies. In this work, we propose to approximate LRM's emerging hierarchical reasoning dynamics as a trajectory within a Finite State Machine (FSM) transitioning among six abstract cognitive states. We demonstrate that these states and transitions can be captured in the latent state of the model. We believe that this representation can have different applications in the interpretability and optimization of LRM models. For example, by analyzing the topology of these transitions, we identify statistical shifts in reasoning strategies that help identify effective reasoning chains from those that fail. To illustrate these potential advantages, we propose Q-Value guided steering, a training-free inference-time control method that treats reasoning as a planning problem. We estimate the long-horizon utility of state transitions and apply sparse, orthogonal activation steering at sentence boundaries to align the CoT generation with optimal reasoning policies. Experiments across four benchmarks (AIME25, MATH-500, GSM8k, and GPQA Diamond) using three stateof-the-art open reasoning models demonstrate that Q-Value steering policy achieves significant performance gains with "surgical" efficiency, often requiring 25× fewer interventions than greedy and weighted baselines, which suggests that reasoning can be effectively controlled by guiding high-level cognitive dynamics rather than micromanaging token generation. Code is available at: https://github.com/shahariar-shibli/CoT-FSM.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Graph of Thoughts: Solving Elaborate Problems with Large Language ModelsMaciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger 等AAAI 2024 · 被引用 1,292 次
- Refusal in Language Models Is Mediated by a Single DirectionAndy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka 等NeurIPS 2024 · 被引用 1,166 次
- RL Squeezes, SFT Expands: A Comparative Study of Reasoning LLMsKohsei Matsutani, Shota Takashiro, Gouki Minegishi, Takeshi Kojima 等ICLR 2026 · 被引用 36 次
- Topology of Reasoning: Understanding Large Reasoning Models through Reasoning Graph PropertiesGouki Minegishi, Hiroki Furuta, Takeshi Kojima, Yusuke Iwasawa 等NeurIPS 2025 · 被引用 31 次
相关 Paper
- Eliciting Chain-of-Thought in Base LLMs via Gradient-Based Representation OptimizationZijian Wang, Yanxiang Ma, Chang XuAAAI 2026
- LLM Reasoning as Trajectories: Step-Specific Representation Geometry and Correctness SignalsLihao Sun, Hang Dong, Bo Qiao, Qingwei Lin 等ACL 2026 · 被引用 9 次
- DenseSteer: Steering Small Language Models towards Dense Math ReasoningYang Ouyang, Shuhang Lin, Jung-Eun KimICML 2026 · 被引用 1 次
- Unlocking the Black Box of Latent Reasoning: An Interpretability-Guided Approach to InterventionShuochen Chang, Tong Bai, Xiaofeng Zhang, Qianli Ma 等ACL 2026 · 被引用 1 次
- Feature Extraction and Steering for Enhanced Chain-of-Thought Reasoning in Language ModelsZihao Li, Xu Wang, Yuzhe Yang, Ziyu Yao 等EMNLP 2025 · 被引用 14 次
