Through the Valley: Path to Effective Long CoT Training for Small Language Models
Renjie Luo, Jiaxi Li, Chen Huang, Wei Lu
Abstract
Long chain-of-thought (CoT) supervision has become a common strategy to enhance reasoning in language models. While effective for large models, we identify a phenomenon we call Long CoT Degradation, in which small language models (SLMs; ≤3B parameters) trained on limited long CoT data experience significant performance deterioration. Through extensive experiments on the Qwen2.5, LLaMA3 and Gemma3 families, we demonstrate that this degradation is widespread across SLMs. In some settings, models trained on only 8k long CoT examples lose up to 75% of their original performance before fine-tuning. Strikingly, we further observe that for some particularly small models, even training on 220k long CoT examples fails to recover or surpass their original performance prior to fine-tuning. Our analysis attributes this effect to error accumulation: while longer responses increase the capacity for multi-step reasoning, they also amplify the risk of compounding mistakes. Furthermore, we find that Long CoT Degradation may negatively impacts downstream reinforcement learning (RL), although this can be alleviated by sufficiently scaled supervised fine-tuning (SFT). Our findings challenge common assumptions about the benefits of long CoT training for SLMs and offer practical guidance for building more effective small-scale reasoning models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8e2d3c84-0a2e-4c8f-a993-22d2757e3fd0Cited by top-tier papers6
- Executable Counterfactuals: Improving LLMs' Causal Reasoning Through CodeAniket Vashishtha, Qirun Dai, Hongyuan Mei, Amit Sharma et al.ICLR 2026 · 11 citations
- How to Fine-Tune a Reasoning Model? A Teacher–Student Cooperation Framework to Synthesize Student-Consistent SFT DataZixian Huang, Kaichen Yang, Xu Huang, Feiyang Hao et al.ICML 2026 · 5 citations
- Which Reasoning Trajectories Teach Students to Reason Better? A Simple Metric of Informative AlignmentYuming Yang, Mingyoung Lai, Wanxu Zhao, Xiaoran Fan et al.ACL 2026 · 5 citations
- VCORE: Variance-Controlled Optimization-based Reweighting for Chain-of-Thought SupervisionXuan Gong, Senmiao Wang, Hanbo Huang, Ruoyu Sun et al.ACL 2026
- MACoT: Synthesizing Chains of Thought for Small Models via Multi-Agent CollaborationGuokai Tang, Feng ZhaoAAAI 2026
Builds on10
- Deep Double Descent: Where Bigger Models and More Data HurtPreetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang et al.ICLR 2020 · 1,108 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base ModelJingcheng Hu, Yinmin Zhang, Qi Han, Daxin Jiang et al.NeurIPS 2025 · 533 citations
- On-Policy Distillation of Language Models: Learning from Self-Generated MistakesRishabh Agarwal, Nino Vieillard, Yongchao Zhou, Piotr Stanczyk et al.ICLR 2024 · 311 citations
- When More is Less: Understanding Chain-of-Thought Length in LLMsYuyang Wu, Yifei Wang, Ziyu Ye, Tianqi Du et al.ICLR 2026 · 225 citations
Related papers
- Large Language Models Are Reasoning TeachersNamgyu Ho, Laura Schmid, Se-Young YunACL 2023 · 102 citations
- Demystifying Long Chain-of-Thought ReasoningShiming Yang, Yuxuan Tong, Xinyao Niu, Graham Neubig et al.ICML 2025
- Towards Efficient Large Language Reasoning Models via Extreme-Ratio Chain-of-Thought CompressionYuntian Tang, Bohan Jia, Wenxuan Huang, Lianyue Zhang et al.ICML 2026 · 5 citations
- Activation Control for Efficiently Eliciting Long Chain-of-thought Ability of Language ModelsZekai Zhao, Qi Liu, Kun Zhou, Zihan Liu et al.NeurIPS 2025 · 10 citations
- Fine-Tuning Large Vision-Language Models as Decision-Making Agents via Reinforcement LearningSimon Zhai, Hao Bai, Zipeng Lin, Jiayi Pan et al.NeurIPS 2024 · 214 citations
