Demystifying Long Chain-of-Thought Reasoning
Shiming Yang, Yuxuan Tong, Xinyao Niu, Graham Neubig, Xiang Yue
Abstract
Scaling inference compute enhances reasoning in large language models (LLMs), with long chainsof-thought (CoTs) enabling strategies like backtracking and error correction. Reinforcement learning (RL) has emerged as a crucial method for developing these capabilities, yet the conditions under which long CoTs emerge remain unclear, and RL training requires careful design choices. In this study, we systematically investigate the mechanics of long CoT reasoning, identifying the key factors that enable models to generate long CoT trajectories. Through extensive supervised fine-tuning (SFT) and RL experiments, we present four main findings: (1) While SFT is not strictly necessary, it simplifies training and improves efficiency; (2) Reasoning capabilities tend to emerge with increased training compute, but their development is not guaranteed, making reward shaping crucial for stabilizing CoT length growth; (3) Scaling verifiable reward signals is critical for RL. We find that leveraging noisy, web-extracted solutions with filtering mechanisms shows strong potential, particularly for out-of-distribution (OOD) tasks such as STEM reasoning; and (4) Core abilities like error correction are inherently present in base models, but incentivizing these skills effectively for complex tasks via RL demands significant compute, and measuring their emergence requires a nuanced approach. These insights provide practical guidance for optimizing training strategies to enhance long CoT reasoning in LLMs. Our code is available at: https://github.com/eddycmu/demystify-long-cot .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b8af0cde-a632-4034-ba86-42ad15358ac9Cited by top-tier papers6
- Better, Faster: Harnessing Self-Improvement in Large Reasoning ModelsQihuang Zhong, Liang Ding, Juhua Liu, Bo Du et al.ICML 2026 · 3 citations
- Contrastive Reasoning Alignment: Reinforcement Learning from Hidden RepresentationsHaozheng Luo, Yimin Wang, Jiahao Yu, Binghui Wang et al.ICML 2026 · 2 citations
- Reasoning Fails Where Step Flow BreaksXiaoyu Xu, Yulan Pan, Xiaosong Yuan, Zhihong Shen et al.ACL 2026 · 1 citation
- Learn to Think: Improving Multimodal Reasoning through Vision-Aware Self-Improvement TrainingQihuang Zhong, Liang Ding, Wenjie Xuan, Juhua Liu et al.ICML 2026 · 1 citation
- Measuring and Mitigating Post-Hoc Rationalization in Reverse Chain-of-Thought GenerationGuangyue Peng, Zongchao Chen, Wen Luo, Yuntao Wen et al.ICML 2026
Builds on13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao et al.ICLR 2024 · 2,082 citations
Related papers
- Provable Benefits of RLVR over SFT for Reasoning Models: Learning to Backtrack EfficientlyStanley Wei, Juno KimICML 2026
- When More is Less: Understanding Chain-of-Thought Length in LLMsYuyang Wu, Yifei Wang, Ziyu Ye, Tianqi Du et al.ICLR 2026 · 225 citations
- Activation Control for Efficiently Eliciting Long Chain-of-thought Ability of Language ModelsZekai Zhao, Qi Liu, Kun Zhou, Zihan Liu et al.NeurIPS 2025 · 10 citations
- Rectifying LLM Thought from Lens of OptimizationJunnan Liu, Hongwei Liu, Songyang Zhang, Kai ChenICLR 2026 · 3 citations
- Incentivizing Reasoning for Advanced Instruction-Following of Large Language ModelsYulei Qin, Gang Li, Zongyi Li, Zihan Xu et al.NeurIPS 2025 · 17 citations
