Learning When to Think: Shaping Adaptive Reasoning in R1-Style Models via Multi-Stage RL
Songjun Tu, Jiahao Lin, Qichao Zhang, Xiangyu Tian, Linjing Li, Xiangyuan Lan, Dongbin Zhao
摘要
Large reasoning models (LRMs) are proficient at generating explicit, step-by-step reasoning sequences before producing final answers. However, such detailed reasoning can introduce substantial computational overhead and latency, particularly for simple problems. To address this over-thinking problem, we explore how to equip LRMs with adaptive thinking capabilities: enabling them to dynamically decide whether or not to engage in explicit reasoning based on problem complexity. Building on R1-style distilled models, we observe that inserting a simple ellipsis ("...") into the prompt can stochastically trigger either a thinking or no-thinking mode, revealing a latent controllability in the reasoning behavior. Leveraging this property, we propose AutoThink, a multi-stage reinforcement learning (RL) framework that progressively optimizes reasoning policies via stage-wise reward shaping. AutoThink learns to invoke explicit reasoning only when necessary, while defaulting to succinct responses for simpler tasks. Experiments on five mainstream mathematical benchmarks demonstrate that AutoThink achieves favorable accuracy-efficiency trade-offs compared to recent prompting and RL-based pruning methods. It can be seamlessly integrated into any R1-style model, including both distilled and further fine-tuned variants. Notably, AutoThink improves relative accuracy by 6.4 percent while reducing token usage by 52 percent on DeepSeek-R1-Distill-Qwen-1.5B, establishing a scalable and adaptive reasoning paradigm for LRMs. Project Page: https://github.com/ScienceOne-AI/AutoThink.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Dynamic Early Exit in Reasoning ModelsChenxu Yang, Qingyi Si, Yongjie Duan, Zheliang Zhu 等ICLR 2026 · 被引用 250 次
- ARES: Multimodal Adaptive Reasoning via Difficulty-Aware Token-Level Entropy ShapingShuang Chen, Hangyu Guo, Yimeng Ye, Shijue Huang 等ICLR 2026 · 被引用 23 次
- R-4B: Incentivizing General-Purpose Auto-Thinking Capability in MLLMs via Bi-Mode Annealing and Reinforce LearningQi Yang, Bolin Ni, Shiming Xiang, Houwen PengCVPR 2026 · 被引用 19 次
- MIRAGE: Towards AI-Generated Image Detection in the WildOucheng Huang, Manxi Lin, Jiexiang Tan, Xiaoxiong Du 等AAAI 2026 · 被引用 6 次
- Audio-Thinker: Guiding Large Audio Language Model When and How to Think via Reinforcement LearningShu Wu, Chenxing Li, Wenfu Wang, Hao Zhang 等AAAI 2026 · 被引用 4 次
它引用的顶会 Paper15
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan 等NeurIPS 2025 · 被引用 2,828 次
- Dynamic Early Exit in Reasoning ModelsChenxu Yang, Qingyi Si, Yongjie Duan, Zheliang Zhu 等ICLR 2026 · 被引用 250 次
- CoT-Valve: Length-Compressible Chain-of-Thought TuningXinyin Ma, Guangnian Wan, Runpeng Yu, Gongfan Fang 等ACL 2025 · 被引用 162 次
- SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for ReasoningYuqian Fu, Tinghong Chen, Jiajun Chai, Xihuai Wang 等ICLR 2026 · 被引用 97 次
相关 Paper
- AdaptThink: Reasoning Models Can Learn When to ThinkJiajie Zhang, Nianyi Lin, Lei Hou, Ling Feng 等EMNLP 2025 · 被引用 3 次
- Thinking-Based Non-Thinking: Solving the Reward Hacking Problem in Training Hybrid Reasoning Models via Reinforcement LearningSiyuan Gan, Jiaheng Liu, Boyan Wang, Tianpei Yang 等ACL 2026
- LEASH: Adaptive Length Penalty and Reward Shaping for Efficient Large Reasoning ModelYanhao Li, Lu Ma, Jiaran Zhang, Lexiang Tang 等ACL 2026 · 被引用 8 次
- Stop Unnecessary Reflection: Training LRMs for Efficient Reasoning with Adaptive Reflection and Length Coordinated PenaltyZewei Yu, Lirong Gao, Yuke Zhu, Bo Zheng 等ICLR 2026 · 被引用 2 次
- Your Models Have Thought Enough: Training Large Reasoning Models to Stop OverthinkingJinyi Han, Ying Huang, Ying Liao, Haiquan Zhao 等ICLR 2026 · 被引用 11 次
