Deliberate Reasoning in Language Models as Structure-Aware Planning with an Accurate World Model
Siheng Xiong, Ali Payani, Yuan Yang, Faramarz Fekri
Abstract
Enhancing the reasoning capabilities of language models (LMs) remains a key challenge, especially for tasks that require complex, multistep decision-making where existing Chain-of-Thought (CoT) approaches struggle with consistency and verification. In this paper, we propose a novel reasoning framework, referred to as Structure-aware Planning with an Accurate World Model (SWAP), that integrates structured knowledge representation with learned planning. Unlike prior methods that rely purely on natural language reasoning, SWAP leverages entailment graphs to encode structured dependencies and enable symbolic verification of intermediate steps. To systematically construct and update the graph, SWAP employs a policy model to propose candidate expansions and a world model to predict structural updates. To improve accuracy, the world model generates multiple alternative updates, and a discriminator re-ranks them based on plausibility. To encourage diverse exploration, we introduce Diversity-based Modelling (DM), which samples candidates from the remaining probability mass after removing previously sampled candidates from the original policy distribution. Additionally, SWAP improves the discrimination accuracy through Contrastive Ranking (CR), which directly compares candidates within prompts and incorporates metaknowledge to improve ranking quality. We evaluate SWAP across diverse reasoning-intensive benchmarks including math reasoning, logical reasoning, and coding tasks. Extensive experiments demonstrate that SWAP significantly improves upon the base models and consistently outperforms existing reasoning methods 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b3910218-c737-41fc-8023-e2f438148bdfCited by top-tier papers7
- How Instruction and Reasoning Data shape Post-Training: Data Quality through the Lens of Layer-wise GradientsMing Li, Yanhong Li, Ziyue Li, Tianyi ZhouACL 2026 · 9 citations
- ChinaTravel: An Open-Ended Travel Planning Benchmark with Compositional Constraint Validation for Language AgentsJie-Jing Shao, Bo-Wen Zhang, Xiao-Wen Yang, Baizhi Chen et al.ICLR 2026 · 4 citations
- Pushing the Boundaries of Natural Reasoning: Interleaved Bonus from Formal-Logic VerificationChuxue Cao, Jinluan Yang, Haoran Li, Kunhao Pan et al.ICML 2026 · 3 citations
- Resisting Manipulative Bots in Meme Coin Copy Trading: A Multi-Agent Approach with Chain-of-Thought ReasoningYichen Luo, Yebo Feng, Jiahua Xu, Yang LiuWWW 2026 · 1 citation
- FlowMAP: Flow Matching for Generalizable Agent PlanningJiarun Fu, Lizhong Ding, Ye Yuan, Qiuning Wei et al.ICML 2026
Builds on21
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan et al.NeurIPS 2023 · 4,972 citations
Related papers
- Reasoning with Language Model is Planning with World ModelShibo Hao, Yi Gu, Haodi Ma, Joshua Jiahua Hong et al.EMNLP 2023 · 109 citations
- Structure Guided Prompt: Instructing Large Language Model in Multi-Step Reasoning by Exploring Graph Structure of the TextKewei Cheng, Nesreen K. Ahmed, Theodore L. Willke, Yizhou SunEMNLP 2024 · 6 citations
- RAS: Retrieval-And-Structuring for Knowledge-Intensive LLM GenerationPengcheng Jiang, Lang Cao, Ruike Zhu, Minhao Jiang et al.ICLR 2026 · 20 citations
- Plan, Verify and Switch: Integrated Reasoning with Diverse X-of-ThoughtsTengxiao Liu, Qipeng Guo, Yuqing Yang, Xiangkun Hu et al.EMNLP 2023 · 7 citations
- Structured Reasoning for LLMs: A Unified Framework for Efficiency and ExplainabilityYubo Dong, Hehe Fan, Linchao Zhu, Yi YangICLR 2026
