The Quality-Utility Paradox: Why High-Reward Data Impairs Small Model Mathematical Reasoning
Haolong Qian, Xianliang Yang, Ma yinuo, Lirong Che, Feng Lu, Ye Guo, Lei Song, Jiang Bian, Chun Yuan
Abstract
Knowledge distillation from powerful reasoning models is widely used to improve Small Language Models (SLMs) on mathematical reasoning, often assuming that traces with higher reward model scores provide more useful supervision. We identify a counterintuitive Quality-Utility Paradox in mathematical reasoning distillation. Data refined or synthesized by a stronger Oracle obtains higher perceived quality according to reward models, yet consistently underperforms traces generated by the SLM itself and selected through rejection sampling across Qwen2.5, LLaMA-3, and DeepSeek families. Our analysis shows that Oracle refinement couples logical repair with distributional drift away from the SLM's native reasoning distribution. This drift increases the learner's adaptation cost and can outweigh the benefit of improved reasoning logic. To test this mechanism, we introduce Style-Aligned Refinement, which preserves the native trajectory of the SLM while retaining logical repair from the Oracle. This intervention lowers adaptation cost and restores downstream utility, allowing distilled SLMs to match or surpass baselines generated by the SLMs themselves. These findings suggest that effective mathematical reasoning distillation should optimize perceived quality together with compatibility between learner and data. The datasets and code are available at https://github.com/Dracoqhl/Quality-Utility-Paradox.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 557ef1a8-4d58-4a58-a543-77e5f58fc7c6Builds on12
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer et al.NeurIPS 2022 · 2,039 citations
- Self-Instruct: Aligning Language Models with Self-Generated InstructionsYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu et al.ACL 2023 · 540 citations
- Self-Play Fine-Tuning Converts Weak Language Models to Strong Language ModelsZixiang Chen, Yihe Deng, Huizhuo Yuan, Kaixuan Ji et al.ICML 2024 · 527 citations
- On-Policy Distillation of Language Models: Learning from Self-Generated MistakesRishabh Agarwal, Nino Vieillard, Yongchao Zhou, Piotr Stanczyk et al.ICLR 2024 · 311 citations
Related papers
- Generative Adversarial Reasoner: Enhancing LLM Reasoning with Adversarial Reinforcement LearningQihao Liu, Luoxin Ye, Wufei Ma, Yu-Cheng Chou et al.ICLR 2026 · 5 citations
- Harnessing Negative Signals: Reinforcement Distillation from Teacher Data for LLM ReasoningShuyao Xu, Cheng Peng, Jiangxuan Long, Weidi Xu et al.ACL 2026 · 3 citations
- Smaller, Weaker, Yet Better: Training LLM Reasoners via Compute-Optimal SamplingHritik Bansal, Arian Hosseini, Rishabh Agarwal, Vinh Q. Tran et al.ICLR 2025 · 1 citation
- Making Expert Reasoning Learnable with Self-DistillationEthan Mendes, Jungsoo Park, Alan RitterICML 2026 · 1 citation
- QCRD: Quality-guided Contrastive Rationale Distillation for Large Language ModelsWei Wang, Zhaowei Li, Qi Xu, Yiqing Cai et al.EMNLP 2025 · 1 citation
