When Simple Problems Wear Complex Costumes: Improving Efficiency in LRM's Adaptive Reasoning
Junnan Ren, Yan Zhang, Qian Chen, Yunhang Shen, Ke Li, Shengchuan Zhang, Liujuan Cao, Rongrong Ji
Abstract
Recent Large Reasoning Models (LRMs) have demonstrated powerful multi-step problem-solving capabilities but often suffer from inefficiency due to an ``overthinking phenomenon", where they apply complex reasoning to simple tasks, resulting in unnecessary computational cost and latency. While adaptive reasoning models that can switch between generating explicit reasoning and producing direct answers offer a potential solution, their effectiveness is compromised by a critical flaw: they are often misled by superficial linguistic complexity, mistaking verbosely phrased simple problems for complex ones. To address this, we propose a two-stage training framework to create a more robust adaptive reasoner. The first stage uses supervised fine-tuning with augmented data—presenting simple problems in both concise and redundant forms—to teach the model to ignore superficial verbosity. Subsequently, a reinforcement learning phase utilizes Group Relative Policy Optimization (GRPO) with a custom reward function to refine the model's adaptive policy, ensuring it selects a reasoning mode based on true task complexity rather than surface-level cues. The resulting model reduces computational overhead without sacrificing accuracy and demonstrates improved robustness to misleading linguistic cues.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on16
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan et al.NeurIPS 2025 · 2,828 citations
Related papers
- Incentivizing Dual Process Thinking for Efficient Large Language Model ReasoningXiaoxue Cheng, Junyi Li, Zhenduo Zhang, Xinyu Tang et al.NeurIPS 2025 · 25 citations
- Think Only When You Need with Large Hybrid-Reasoning ModelsLingjie Jiang, Xun Wu, Shaohan Huang, Qingxiu Dong et al.NeurIPS 2025 · 71 citations
- Efficiently Learning To Reason or Not to Reason: Root-token Policy Optimization for Adaptive ThinkingTaehyeon Kim, Hyunsoo Lee, Youngsoo Jang, Moontae LeeACL 2026
- ARM: Adaptive Reasoning ModelSiye Wu, Jian Xie, Yikai Zhang, Aili Chen et al.NeurIPS 2025 · 31 citations
- MARS: Multimodal Adaptive Reasoning Model for Avoiding OverthinkingTan Yue, Qiong Wu, Dongyan ZhaoAAAI 2026
