Learning When to Think: Shaping Adaptive Reasoning in R1-Style Models via Multi-Stage RL
Songjun Tu, Jiahao Lin, Qichao Zhang, Xiangyu Tian, Linjing Li, Xiangyuan Lan, Dongbin Zhao
Abstract
Large reasoning models (LRMs) are proficient at generating explicit, step-by-step reasoning sequences before producing final answers. However, such detailed reasoning can introduce substantial computational overhead and latency, particularly for simple problems. To address this over-thinking problem, we explore how to equip LRMs with adaptive thinking capabilities: enabling them to dynamically decide whether or not to engage in explicit reasoning based on problem complexity. Building on R1-style distilled models, we observe that inserting a simple ellipsis ("...") into the prompt can stochastically trigger either a thinking or no-thinking mode, revealing a latent controllability in the reasoning behavior. Leveraging this property, we propose AutoThink, a multi-stage reinforcement learning (RL) framework that progressively optimizes reasoning policies via stage-wise reward shaping. AutoThink learns to invoke explicit reasoning only when necessary, while defaulting to succinct responses for simpler tasks. Experiments on five mainstream mathematical benchmarks demonstrate that AutoThink achieves favorable accuracy-efficiency trade-offs compared to recent prompting and RL-based pruning methods. It can be seamlessly integrated into any R1-style model, including both distilled and further fine-tuned variants. Notably, AutoThink improves relative accuracy by 6.4 percent while reducing token usage by 52 percent on DeepSeek-R1-Distill-Qwen-1.5B, establishing a scalable and adaptive reasoning paradigm for LRMs. Project Page: https://github.com/ScienceOne-AI/AutoThink.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8c243ce2-fdf5-4229-9ac5-7a71ef9b78dbCited by top-tier papers14
- Dynamic Early Exit in Reasoning ModelsChenxu Yang, Qingyi Si, Yongjie Duan, Zheliang Zhu et al.ICLR 2026 · 250 citations
- ARES: Multimodal Adaptive Reasoning via Difficulty-Aware Token-Level Entropy ShapingShuang Chen, Hangyu Guo, Yimeng Ye, Shijue Huang et al.ICLR 2026 · 23 citations
- R-4B: Incentivizing General-Purpose Auto-Thinking Capability in MLLMs via Bi-Mode Annealing and Reinforce LearningQi Yang, Bolin Ni, Shiming Xiang, Houwen PengCVPR 2026 · 19 citations
- MIRAGE: Towards AI-Generated Image Detection in the WildOucheng Huang, Manxi Lin, Jiexiang Tan, Xiaoxiong Du et al.AAAI 2026 · 6 citations
- Audio-Thinker: Guiding Large Audio Language Model When and How to Think via Reinforcement LearningShu Wu, Chenxing Li, Wenfu Wang, Hao Zhang et al.AAAI 2026 · 4 citations
Builds on15
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan et al.NeurIPS 2025 · 2,828 citations
- Dynamic Early Exit in Reasoning ModelsChenxu Yang, Qingyi Si, Yongjie Duan, Zheliang Zhu et al.ICLR 2026 · 250 citations
- CoT-Valve: Length-Compressible Chain-of-Thought TuningXinyin Ma, Guangnian Wan, Runpeng Yu, Gongfan Fang et al.ACL 2025 · 162 citations
- SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for ReasoningYuqian Fu, Tinghong Chen, Jiajun Chai, Xihuai Wang et al.ICLR 2026 · 97 citations
Related papers
- AdaptThink: Reasoning Models Can Learn When to ThinkJiajie Zhang, Nianyi Lin, Lei Hou, Ling Feng et al.EMNLP 2025 · 3 citations
- Thinking-Based Non-Thinking: Solving the Reward Hacking Problem in Training Hybrid Reasoning Models via Reinforcement LearningSiyuan Gan, Jiaheng Liu, Boyan Wang, Tianpei Yang et al.ACL 2026
- LEASH: Adaptive Length Penalty and Reward Shaping for Efficient Large Reasoning ModelYanhao Li, Lu Ma, Jiaran Zhang, Lexiang Tang et al.ACL 2026 · 8 citations
- Stop Unnecessary Reflection: Training LRMs for Efficient Reasoning with Adaptive Reflection and Length Coordinated PenaltyZewei Yu, Lirong Gao, Yuke Zhu, Bo Zheng et al.ICLR 2026 · 2 citations
- Your Models Have Thought Enough: Training Large Reasoning Models to Stop OverthinkingJinyi Han, Ying Huang, Ying Liao, Haiquan Zhao et al.ICLR 2026 · 11 citations
