Can We Further Elicit Reasoning in LLMs? Critic-Guided Planning with Retrieval-Augmentation for Solving Challenging Tasks
Xingxuan Li, Weiwen Xu, Ruochen Zhao, Fangkai Jiao, Shafiq Joty, Lidong Bing
摘要
State-of-the-art large language models (LLMs) exhibit impressive problemsolving capabilities but may struggle with complex reasoning and factual correctness. Existing methods harness the strengths of chain-of-thought (CoT) and retrieval-augmented generation (RAG) to decompose a complex problem into simpler steps and apply retrieval to improve factual correctness. These methods work well on straightforward reasoning tasks but often falter on challenging tasks such as competitive programming and mathematics, due to frequent reasoning errors and irrelevant knowledge retrieval. To address this, we introduce Critic-guided planning with Retrieval-augmentation, CR-Planner, a novel framework that leverages fine-tuned critic models to guide both reasoning and retrieval processes through planning. CR-Planner solves a problem by iteratively selecting and executing sub-goals. Initially, it identifies the most promising sub-goal from reasoning, query generation, and retrieval, guided by rewards given by a critic model named sub-goal critic. It then executes this sub-goal through sampling and selecting the optimal output based on evaluations from another critic model named execution critic. This iterative process, informed by retrieved information and critic models, enables CR-Planner to effectively navigate the solution space towards the final answer. We employ Monte Carlo Tree Search (MCTS) to collect the data for training the critic models, allowing for a systematic exploration of action sequences and their long-term impacts. We validate CR-Planner on challenging domain-knowledge-intensive and reasoning-heavy tasks, including competitive programming, theorem-driven math reasoning, and complex domain retrieval problems. Our experiments demonstrate that CR-Planner significantly outperforms baselines, highlighting its effectiveness in addressing challenging problems by improving both reasoning and retrieval. 1 * Xingxuan Li is under the Joint Ph.D. Program between DAMO Academy and Nanyang Technological University. 1 We will make our code and data publicly available.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- ReARTeR: Retrieval-Augmented Reasoning with Trustworthy Process RewardingZhongxiang Sun, Qipeng Wang, Weijie Yu, Xiaoxue Zang 等SIGIR 2025 · 被引用 5 次
- PosterMate: Audience-driven Collaborative Persona Agents for Poster DesignDonghoon Shin, Daniel Lee, Gary Hsieh, Gromit Yeuk-Yin ChanUIST 2025 · 被引用 3 次
- CP-Search: A Chain Progressive Search Training Framework Incentivizing the Cognitive Behaviors for Searching in LLMsZehua Wang, Shipeng Li, Buzhou TangAAAI 2026
- REAP: Enhancing RAG with Recursive Evaluation and Adaptive Planning for Multi-Hop Question AnsweringYijie Zhu, Haojie Zhou, Wanting Hong, Tailin Liu 等AAAI 2026
- A Survey of Reasoning-Intensive Retrieval: Progress and ChallengesYiyang Wei, Tingyu Song, Siyue Zhang, Yilun ZhaoACL 2026
它引用的顶会 Paper15
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan 等NeurIPS 2023 · 被引用 5,828 次
相关 Paper
- Retrieval is Not Enough: Enhancing RAG through Test-Time Critique and OptimizationJiaqi Wei, Hao Zhou, Xiang Zhang, Di Zhang 等NeurIPS 2025 · 被引用 14 次
- OPERA: A Reinforcement Learning-Enhanced Orchestrated Planner-Executor Architecture for Reasoning-Oriented Multi-Hop RetrievalYu Liu, Yanbing Liu, Fangfang Yuan, Cong Cao 等AAAI 2026 · 被引用 4 次
- InstructRAG: Leveraging Retrieval-Augmented Generation on Instruction Graphs for LLM-Based Task PlanningZheng Wang, Shu Xian Teo, Jun Jie Chew, Wei ShiSIGIR 2025 · 被引用 4 次
- Plan Then Action: High-Level Planning Guidance Reinforcement Learning for LLM ReasoningZhihao Dou, Qinjian Zhao, Zhongwei Wan, Zhang Dinggen 等ICML 2026 · 被引用 24 次
- AtomR: Atomic Operator-Empowered Large Language Models for Heterogeneous Knowledge ReasoningAmy Xin, Jinxin Liu, Zijun Yao, Zhicheng Lee 等KDD 2025 · 被引用 1 次
