Dissecting Failure Dynamics in Large Language Model Reasoning
Wei Zhu, Jian Zhang, Lixing Yu, Kun Yue, Zhiwen Tang
Abstract
Large Language Models (LLMs) achieve strong performance through extended inference-time deliberation, yet how their reasoning failures arise remains poorly understood. By analyzing model-generated reasoning trajectories, we find that errors are not uniformly distributed but often originate from a small number of early transition points, after which reasoning remains locally coherent but globally incorrect. These transitions coincide with localized spikes in token-level entropy, and alternative continuations from the same intermediate state can still lead to correct solutions. Based on these observations, we introduce GUARD 1 , a targeted inference-time framework that probes and redirects critical transitions using uncertainty signals. Empirical evaluations across multiple benchmarks confirm that interventions guided by these failure dynamics lead to more reliable reasoning outcomes. Our findings highlight the importance of understanding when and how reasoning first deviates, complementing existing approaches that focus on scaling inference-time computation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 23abe19b-176b-4f36-b45c-b7ded03546aaCited by top-tier papers2
- AttnPO: Attention-Guided Process Supervision for Efficient ReasoningShuaiyi Nie, Siyu Ding, Wenyuan Zhang, Linhao Yu et al.ACL 2026 · 21 citations
- AIM-CoT: Active Information-driven Multimodal Chain-of-Thought for Vision-Language ReasoningXiping Li, Jianghong MaACL 2026 · 3 citations
Builds on20
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le et al.ICLR 2023 · 681 citations
- Dynamic Early Exit in Reasoning ModelsChenxu Yang, Qingyi Si, Yongjie Duan, Zheliang Zhu et al.ICLR 2026 · 250 citations
Related papers
- ReasoningGuard: Safeguarding Large Reasoning Models with Inference-time Safety Aha MomentsYuquan Wang, Mi Zhang, Yining Wang, Geng Hong et al.ACL 2026 · 2 citations
- Characterizing and Mitigating Reasoning Drift in Large Language ModelsYufeng Zhang, Xuepeng Wang, Lingxiang Wu, Jinqiao WangICLR 2026
- LLM Reasoning as Trajectories: Step-Specific Representation Geometry and Correctness SignalsLihao Sun, Hang Dong, Bo Qiao, Qingwei Lin et al.ACL 2026 · 9 citations
- Intervene When It Doubts: Conjunction-Guided Interactive ReasoningQianyue Wang, Jinwu Hu, Yaofo Chen, Yufeng Wang et al.ICML 2026
- Scaling Reasoning Hop Exposes Weaknesses: Demystifying and Improving Hop Generalization in Large Language ModelsZhaoyi Li, Jiatong Li, Gangwei Jiang, Linqi Song et al.ICLR 2026 · 1 citation
