Subtle Errors in Reasoning: Preference Learning via Error-injected Self-editing
Kaishuai Xu, Tiezheng Yu, Wenjun Hou, Yi Cheng, Chak Tou Leong, Liangyou Li, Xin Jiang, Lifeng Shang, Qun Liu, Wenjie Li
Abstract
Large Language Models (LLMs) have exhibited strong mathematical reasoning prowess, tackling tasks ranging from basic arithmetic to advanced competition-level problems. However, frequently occurring subtle yet critical errors, such as miscalculations or incorrect substitutions, limit the LLMs' full potential. Existing studies to improve mathematical ability typically involve applying preference learning to step-wise solution pairs. Although these methods leverage samples of varying granularity to mitigate reasoning errors, they overlook critical subtle errors. In this work, we propose a novel preference learning framework called eRror-Injected Self-Editing (RISE), which injects predefined subtle errors into pivotal tokens in reasoning or computation steps to construct hard pairs for error mitigation. In detail, RISE uses the LLM itself to edit a small number of tokens in the solution, injecting designed subtle errors. Then, pairs composed of self-edited solutions and their corresponding correct ones, along with pairs of correct and incorrect solutions obtained through sampling, are used together for subtle error-aware DPO training. Compared with other preference learning methods, RISE further refines the training objective without requiring fine-grained sampling or preference annotation. Extensive experiments validate the effectiveness of RISE, with preference learning on Qwen2-7B-Instruct yielding notable improvements of 3.0% on GSM8K and 7.9% on MATH with only 4.5K training samples. Moreover, the effect of error mitigation extends from mathematical reasoning to logical reasoning and code generation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on12
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- MetaMath: Bootstrap Your Own Mathematical Questions for Large Language ModelsLonghui Yu, Weisen Jiang, Han Shi, Jincheng Yu et al.ICLR 2024 · 637 citations
- MAmmoTH: Building Math Generalist Models through Hybrid Instruction TuningXiang Yue, Xingwei Qu, Ge Zhang, Yao Fu et al.ICLR 2024 · 558 citations
Related papers
- Critical Tokens Matter: Token-Level Contrastive Estimation Enhances LLM's Reasoning CapabilityZicheng Lin, Tian Liang, Jiahao Xu, Qiuzhi Liu et al.ICML 2025
- Self-Error-Instruct: Generalizing from Errors for LLMs Mathematical ReasoningErxin Yu, Jing Li, Ming Liao, Qi Zhu et al.ACL 2025 · 4 citations
- S^3cMath: Spontaneous Step-Level Self-Correction Makes Large Language Models Better Mathematical ReasonersYuchen Yan, Jin Jiang, Yang Liu, Yixin Cao et al.AAAI 2025 · 19 citations
- Self-Training with Direct Preference Optimization Improves Chain-of-Thought ReasoningTianduo Wang, Shichen Li, Wei LuACL 2024
- Uncertainty-Aware Iterative Preference Optimization for Enhanced LLM ReasoningLei Li, Hehuan Liu, Yaxin Zhou, ZhaoYang Gui et al.ACL 2025 · 3 citations
