Learning Math Reasoning from Self-Sampled Correct and Partially-Correct Solutions
Ansong Ni, Jeevana Priya Inala, Chenglong Wang, Alex Polozov, Christopher Meek, Dragomir Radev, Jianfeng Gao
Abstract
Pretrained language models have shown superior performance on many natural language processing tasks, yet they still struggle at multi-step formal reasoning tasks like grade school math problems. One key challenge of finetuning them to solve such math reasoning problems is that many existing datasets only contain one reference solution for each problem, despite the fact that there are often alternative solutions resembling different reasoning paths to the final answer. This way, the finetuned models are biased towards the limited reference solutions, which limits their generalization to unseen examples. To mitigate this issue, we propose to let the model perform sampling during training and learn from both selfsampled fully-correct solutions, which yield the correct answer upon execution, and partially-correct solutions, whose intermediate state matches an intermediate state of a known correct solution. We show that our use of self-sampled correct and partially-correct solutions can benefit learning and help guide the sampling process, leading to more efficient exploration of the solution space. Additionally, we explore various training objectives to support learning from multiple solutions per example and find they greatly affect the performance. Experiments on two math reasoning datasets show the effectiveness of our method compared to learning from a single reference solution with MLE, where we improve PASS@100 from 35.5% to 44.5% for GSM8K, and 27.6% to 36.2% PASS@80 for MathQA. Such improvements are also consistent across different model sizes. Our code is available at https://github.com/microsoft/TraceCodegen . * Majority of the work done during an internship at Microsoft Research. † Work started while at Microsoft Research, paper contribution limited to proof-reading. ‡ Work initiated while at Microsoft.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d67d2693-63b2-424c-b7ac-264c1661a7faCited by top-tier papers16
- Chain of Preference Optimization: Improving Chain-of-Thought Reasoning in LLMsXuan Zhang, Chao Du, Tianyu Pang, Qian Liu et al.NeurIPS 2024 · 177 citations
- Large Language Model Cascades with Mixture of Thought Representations for Cost-Efficient ReasoningMurong Yue, Jie Zhao, Min Zhang, Liang Du et al.ICLR 2024 · 153 citations
- NExT: Teaching Large Language Models to Reason about Code ExecutionAnsong Ni, Miltiadis Allamanis, Arman Cohan, Yinlin Deng et al.ICML 2024 · 73 citations
- Coevolving with the Other You: Fine-Tuning LLM with Sequential Cooperative Multi-Agent Reinforcement LearningHao Ma, Tianyi Hu, Zhiqiang Pu, Boyin Liu et al.NeurIPS 2024 · 54 citations
- Interpreting and Improving Large Language Models in Arithmetic CalculationWei Zhang, Chaoqun Wan, Yonggang Zhang, Yiu-ming Cheung et al.ICML 2024 · 47 citations
Builds on6
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Semantic Evaluation for Text-to-SQL with Distilled Test SuitesRuiqi Zhong, Tao Yu, Dan KleinEMNLP 2020 · 88 citations
- CodeT: Code Generation with Generated TestsBei Chen, Fengji Zhang, Anh Nguyen, Daoguang Zan et al.ICLR 2023 · 64 citations
- Latent Execution for Neural Program Synthesis Beyond Domain-Specific LanguagesXinyun Chen, Dawn Song, Yuandong TianNeurIPS 2021 · 56 citations
- Representing Partial Programs with Blended Abstract SemanticsMaxwell I. Nye, Yewen Pu, Matthew Bowers, Jacob Andreas et al.ICLR 2021 · 23 citations
Related papers
- Learning to Better Search with Language Models via Guided Reinforced Self-TrainingSeungyong Moon, Bumsoo Park, Hyun Oh SongNeurIPS 2025 · 3 citations
- ReFT: Reasoning with Reinforced Fine-TuningLuong Quoc Trung, Xinbo Zhang, Zhanming Jie, Peng Sun et al.ACL 2024
- Physics of Language Models: Part 2.2, How to Learn From Mistakes on Grade-School Math ProblemsTian Ye, Zicheng Xu, Yuanzhi Li, Zeyuan Allen-ZhuICLR 2025 · 2 citations
- S^3cMath: Spontaneous Step-Level Self-Correction Makes Large Language Models Better Mathematical ReasonersYuchen Yan, Jin Jiang, Yang Liu, Yixin Cao et al.AAAI 2025 · 19 citations
- MathSE: Improving Multimodal Mathematical Reasoning via Self-Evolving Iterative Reflection and Reward-Guided Fine-TuningJinhao Chen, Zhen Yang, Jianxin Shi, Tianyu Wo et al.AAAI 2026 · 2 citations
