Structured Reasoning for LLMs: A Unified Framework for Efficiency and Explainability
Yubo Dong, Hehe Fan, Linchao Zhu, Yi Yang
Abstract
Recent Large Language Models (LLMs) have made remarkable progress, but they still struggle with complex reasoning tasks such as logical deduction and planning. This is partly because they rely primarily on token-level probability relationships, which limits their ability to reason effectively.
In this paper, inspired by cognitive science and neurosymbolic AI, we introduce Structured Reasoning, which aimes at enhancing the reasoning capabilities of LLMs from the step level.
To this end, we first collect high‑frequency, domain‑agnostic reasoning step tags and construct a structured reasoning dataset with those tags.
Then, we treat a reasoning process as a directed acyclic graph, where the vertices represent steps and the edges indicate the direction of reasoning.
In this context, an efficient reasoning process corresponds to, or can be characterized by, a sparse reasoning graph.
To construct reasoning graphs, we introduce structured tags for reliable step extraction from LLM outputs. For single-graph optimization, we propose the MaxFlow reward, which rewards graphs with balanced node contributions and fewer redundant steps. The quality of a sparse reasoning graph can be reflected by the total flow from all steps to the final answer. For multi-graph comparison, we propose the LCS reward, which selects reliable reasoning paths by identifying optimal common subsequences (consecutive steps) shared across multiple generated responses (sequences).
Experiments with DeepSeek-R1-Distill-Qwen-1.5B and 7B models show that our method consistently outperforms GRPO and other carefully tuned baselines across various context lengths (0.5k–8k).
Structured Reasoning shows particular strength in efficiency (better performance with fewer steps) and stability (consistently generating high-quality outputs across a temperature range of 0.1 to 1.0).
Methods and examples is currently available on our website: https://cnsdqd-dyb.github.io/structured-reasoning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on8
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan et al.NeurIPS 2025 · 2,828 citations
- Training Language Models to Reason EfficientlyDaman Arora, Andrea ZanetteNeurIPS 2025 · 270 citations
- Can Language Models Learn to Skip Steps?Tengxiao Liu, Qipeng Guo, Xiangkun Hu, Cheng Jiayang et al.NeurIPS 2024 · 92 citations
- Exploring Data Scaling Trends and Effects in Reinforcement Learning from Human FeedbackWei Shen, Guanlin Liu, Yu Yue, Ruofei Zhu et al.NeurIPS 2025 · 33 citations
- OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific ProblemsChaoqun He, Renjie Luo, Yuzhuo Bai, Shengding Hu et al.ACL 2024 · 18 citations
Related papers
- Structure Guided Prompt: Instructing Large Language Model in Multi-Step Reasoning by Exploring Graph Structure of the TextKewei Cheng, Nesreen K. Ahmed, Theodore L. Willke, Yizhou SunEMNLP 2024 · 6 citations
- SEER: Facilitating Structured Reasoning and Explanation via Reinforcement LearningGuoxin Chen, Kexin Tang, Chao Yang, Fuying Ye et al.ACL 2024 · 5 citations
- Generative Adversarial Reasoner: Enhancing LLM Reasoning with Adversarial Reinforcement LearningQihao Liu, Luoxin Ye, Wufei Ma, Yu-Cheng Chou et al.ICLR 2026 · 5 citations
- Learn to Reason Efficiently with Adaptive Length-based Reward ShapingWei Liu, Ruochen Zhou, Yiyun Deng, Yuzhen Huang et al.ICLR 2026 · 88 citations
- Graph Reasoning Paradigm: Structured and Symbolic Reasoning with Topology-Aware Reinforcement Learning for Large Language ModelsRunxuan Liu, Xianhao Ou, Xinyan Ma, Jiyuan Wang et al.ACL 2026
