Strategy Executability in Mathematical Reasoning: Leveraging Human–Model Differences for Effective Guidance
Weida Liang, Yiyou Sun, Shuyuan Nan, Chuang Li, Dawn Song, Kenji Kawaguchi
Abstract
Example-based guidance is widely used to improve mathematical reasoning at inference time, yet its effectiveness is highly unstable across problems and models—even when the guidance is correct and problem-relevant. We show that this instability arises from a previously underexplored gap between strategy usage—whether a reasoning strategy appears in successful solutions—and strategy executability—whether the strategy remains effective when instantiated as guidance for a target model. Through a controlled analysis of paired human-written and model-generated solutions, we identify a systematic dissociation between usage and executability: human- and model-derived strategies differ in structured, domain-dependent ways, leading to complementary strengths and consistent source-dependent reversals under guidance. Building on this diagnosis, we propose Selective Strategy Retrieval (SSR), a test-time framework that explicitly models executability by selectively retrieving and combining strategies using empirical, multi-route, source-aware signals. Across multiple mathematical reasoning benchmarks, SSR yields reliable and consistent improvements over direct solving, in-context learning, and single-source guidance, improving accuracy by up to points on AIME25 and points on Apex for compact reasoning models. Code and benchmark are publicly available at: https://github.com/lwd17/strategy-execute-pipeline.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan et al.NeurIPS 2023 · 4,972 citations
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer et al.NeurIPS 2022 · 2,039 citations
Related papers
- Sticker-TTS: Learn to Utilize Historical Experience with a Sticker-driven Test-Time Scaling FrameworkJie Chen, Jinhao Jiang, Yingqian Min, Zican Dong et al.EMNLP 2025
- Find Tailored Step Example for Next Step: a Targeted Step-wise Retrieval Framework for Guiding LLM ReasoningCheng Yang, Zhenya Huang, Liyang He, Weibo Gao et al.KDD 2026
- Towards Effective Code-Integrated ReasoningFei Bai, Yingqian Min, Beichen Zhang, Zhipeng Chen et al.AAAI 2026
- What If We Allocate Test-Time Compute Adaptively?Ahsan Bilal, Muhammad Ahmed Mohsin, Muhammad Umer, Ali Subhan et al.ICML 2026 · 3 citations
- TRIM: Hybrid Inference via Targeted Stepwise Routing in Multi-Step Reasoning TasksVansh Kapoor, Aman Gupta, Hao Chen, Anurag Beniwal et al.ICLR 2026 · 7 citations
