Deliberate Reasoning in Language Models as Structure-Aware Planning with an Accurate World Model
Siheng Xiong, Ali Payani, Yuan Yang, Faramarz Fekri
摘要
Enhancing the reasoning capabilities of language models (LMs) remains a key challenge, especially for tasks that require complex, multistep decision-making where existing Chain-of-Thought (CoT) approaches struggle with consistency and verification. In this paper, we propose a novel reasoning framework, referred to as Structure-aware Planning with an Accurate World Model (SWAP), that integrates structured knowledge representation with learned planning. Unlike prior methods that rely purely on natural language reasoning, SWAP leverages entailment graphs to encode structured dependencies and enable symbolic verification of intermediate steps. To systematically construct and update the graph, SWAP employs a policy model to propose candidate expansions and a world model to predict structural updates. To improve accuracy, the world model generates multiple alternative updates, and a discriminator re-ranks them based on plausibility. To encourage diverse exploration, we introduce Diversity-based Modelling (DM), which samples candidates from the remaining probability mass after removing previously sampled candidates from the original policy distribution. Additionally, SWAP improves the discrimination accuracy through Contrastive Ranking (CR), which directly compares candidates within prompts and incorporates metaknowledge to improve ranking quality. We evaluate SWAP across diverse reasoning-intensive benchmarks including math reasoning, logical reasoning, and coding tasks. Extensive experiments demonstrate that SWAP significantly improves upon the base models and consistently outperforms existing reasoning methods 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- How Instruction and Reasoning Data shape Post-Training: Data Quality through the Lens of Layer-wise GradientsMing Li, Yanhong Li, Ziyue Li, Tianyi ZhouACL 2026 · 被引用 9 次
- ChinaTravel: An Open-Ended Travel Planning Benchmark with Compositional Constraint Validation for Language AgentsJie-Jing Shao, Bo-Wen Zhang, Xiao-Wen Yang, Baizhi Chen 等ICLR 2026 · 被引用 4 次
- Pushing the Boundaries of Natural Reasoning: Interleaved Bonus from Formal-Logic VerificationChuxue Cao, Jinluan Yang, Haoran Li, Kunhao Pan 等ICML 2026 · 被引用 3 次
- Resisting Manipulative Bots in Meme Coin Copy Trading: A Multi-Agent Approach with Chain-of-Thought ReasoningYichen Luo, Yebo Feng, Jiahua Xu, Yang LiuWWW 2026 · 被引用 1 次
- FlowMAP: Flow Matching for Generalizable Agent PlanningJiarun Fu, Lizhong Ding, Ye Yuan, Qiuning Wei 等ICML 2026
它引用的顶会 Paper21
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan 等NeurIPS 2023 · 被引用 5,828 次
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan 等NeurIPS 2023 · 被引用 4,972 次
相关 Paper
- Reasoning with Language Model is Planning with World ModelShibo Hao, Yi Gu, Haodi Ma, Joshua Jiahua Hong 等EMNLP 2023 · 被引用 109 次
- Structure Guided Prompt: Instructing Large Language Model in Multi-Step Reasoning by Exploring Graph Structure of the TextKewei Cheng, Nesreen K. Ahmed, Theodore L. Willke, Yizhou SunEMNLP 2024 · 被引用 6 次
- RAS: Retrieval-And-Structuring for Knowledge-Intensive LLM GenerationPengcheng Jiang, Lang Cao, Ruike Zhu, Minhao Jiang 等ICLR 2026 · 被引用 20 次
- Plan, Verify and Switch: Integrated Reasoning with Diverse X-of-ThoughtsTengxiao Liu, Qipeng Guo, Yuqing Yang, Xiangkun Hu 等EMNLP 2023 · 被引用 7 次
- Structured Reasoning for LLMs: A Unified Framework for Efficiency and ExplainabilityYubo Dong, Hehe Fan, Linchao Zhu, Yi YangICLR 2026
