Learning from Mistakes: Negative Reasoning Samples Enhance Out-of-Domain Generalization
Tian Xueyun, Minghua Ma, Bingbing Xu, Nuoyan Lyu, Wei Li, Heng Dong, Zheng Chu, Yuanzhuo Wang, Huawei Shen
Abstract
Supervised fine-tuning (SFT) on chain-ofthought (CoT) trajectories demonstrations is a common approach for enabling reasoning in large language models. Standard practices typically only retain trajectories with correct final answers (positives) while ignoring the rest (negatives). We argue that this paradigm discards substantial supervision and exacerbates overfitting, limiting out-of-domain (OOD) generalization. Specifically, we surprisingly find that incorporating negative trajectories into SFT yields substantial OOD generalization gains over positive-only training, as these trajectories often retain valid intermediate reasoning despite incorrect final answers. To understand this effect in depth, we systematically analyze data, training dynamics, and inference behavior, identifying 22 recurring patterns in negative chains that serve a dual role: they moderate loss descent to mitigate overfitting during training and boost policy entropy by 35.67% during inference to facilitate exploration. Motivated by these observations, we further propose Gainbased LOss Weighting (GLOW), an adaptive, sample-aware scheme that exploits such distinctive training dynamics by rescaling per-sample loss based on inter-epoch progress. Empirically, GLOW efficiently leverages unfiltered trajectories, yielding a 5.51% OOD gain over positive-only SFT on Qwen2.5-7B and boosting MMLU from 72.82% to 76.47% as an RL initialization. Code is available at Github 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6d4fbf59-c9ee-42c7-a0e4-cd163ce2baa3Cited by top-tier papers1
Ask how each one uses itBuilds on9
- STaR: Bootstrapping Reasoning With ReasoningEric Zelikman, Yuhuai Wu, Jesse Mu, Noah D. GoodmanNeurIPS 2022 · 1,126 citations
- The Surprising Effectiveness of Negative Reinforcement in LLM ReasoningXinyu Zhu, Mengzhou Xia, Zhepei Wei, Wei-Lin Chen et al.NeurIPS 2025 · 177 citations
- Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM ReasoningMaggie Ziyu Huan, Yuetai Li, Tuney Zheng, Xiaoyu Xu et al.ICML 2026 · 102 citations
- ACPBench: Reasoning About Action, Change, and PlanningHarsha Kokel, Michael Katz, Kavitha Srinivas, Shirin SohrabiAAAI 2025 · 35 citations
- Customizing Language Model Responses with Contrastive In-Context LearningXiang Gao, Kamalika DasAAAI 2024 · 23 citations
Related papers
- VCORE: Variance-Controlled Optimization-based Reweighting for Chain-of-Thought SupervisionXuan Gong, Senmiao Wang, Hanbo Huang, Ruoyu Sun et al.ACL 2026
- When More is Less: Understanding Chain-of-Thought Length in LLMsYuyang Wu, Yifei Wang, Ziyu Ye, Tianqi Du et al.ICLR 2026 · 225 citations
- Demystifying Long Chain-of-Thought ReasoningShiming Yang, Yuxuan Tong, Xinyao Niu, Graham Neubig et al.ICML 2025
- InT: Self-Proposed Interventions Enable Credit Assignment in LLM ReasoningMatthew Y. R. Yang, Hao Bai, Ian Wu, Gene Yang et al.ICLR 2026 · 12 citations
- Through the Valley: Path to Effective Long CoT Training for Small Language ModelsRenjie Luo, Jiaxi Li, Chen Huang, Wei LuEMNLP 2025
