Find Tailored Step Example for Next Step: a Targeted Step-wise Retrieval Framework for Guiding LLM Reasoning
Cheng Yang, Zhenya Huang, Liyang He, Weibo Gao, Wei Huang, Wei Tong, Xin Li
Abstract
Large language models (LLMs) have shown strong performance in mathematical reasoning, supported by approaches such as In-Context Learning (ICL) and Retrieval-Augmented Generation (RAG). However, existing methods often provide problem-level examples, which is too coarse-grained for multi-step reasoning to cause informational redundancy, and structural misalignment. To address this limitation, we propose Step-wise Training for In-context Reasoning (STIR) to provide step-synchronized and logically targeted guidance to enhance the model's mathematical reasoning capabilities. STIR enables a model to dynamically decide when to retrieve a single logically consistent next step, just using the current problem and its intermediate state as the query. First, We decompose expert solutions into Step-Level Reasoning Units inspired by human thinking patterns. Leveraging this data, a Step Retriever is trained for logical continuity to map current reasoning states to relevant subsequent steps. Then a Step Reasoner is trained to decide when to retrieve tailored step examples and incorporates this guidance into reasoning. We further extend STIR with a Process-aware Reinforcement Learning phase using Group Relative Policy Optimization to learn to self-formulate search queries and optimizes the decision-making policy. Experiments on seven benchmarks demonstrate that STIR achieves accuracy improvements ranging from 1.86% to 17.96%, maintaining lower token efficiency than baselines. Analysis via our proposed DSM, TCN and RCR metrics shows that STIR improves reasoning capability, achieving DSM scores ranging from 4.54 to 29.87 across backbones and significant improvements over the baseline in both TCN and RCR.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 53fe99fb-5d9d-464a-a8a3-4b4f6219847cRelated papers
- Plan Then Action: High-Level Planning Guidance Reinforcement Learning for LLM ReasoningZhihao Dou, Qinjian Zhao, Zhongwei Wan, Zhang Dinggen et al.ICML 2026 · 24 citations
- Step Guided Reasoning: Improving Mathematical Reasoning using Guidance Generation and Step ReasoningLang Cao, Yingtian Zou, Chao Peng, Renhong Chen et al.EMNLP 2025 · 7 citations
- Selection-Inference: Exploiting Large Language Models for Interpretable Logical ReasoningAntonia Creswell, Murray Shanahan, Irina HigginsICLR 2023 · 110 citations
- Learning to Reason over Continuous Tokens with Reinforcement LearningYiran Zhao, Yuhui Xu, Doyen Sahoo, Caiming Xiong et al.ICLR 2026 · 1 citation
- ReARTeR: Retrieval-Augmented Reasoning with Trustworthy Process RewardingZhongxiang Sun, Qipeng Wang, Weijie Yu, Xiaoxue Zang et al.SIGIR 2025 · 5 citations
