Incentivizing Dual Process Thinking for Efficient Large Language Model Reasoning
Xiaoxue Cheng, Junyi Li, Zhenduo Zhang, Xinyu Tang, Xin Zhao, Xinyu Kong, Zhiqiang Zhang
Abstract
Large reasoning models (LRMs) have demonstrated strong performance on complex reasoning tasks, but often suffer from overthinking, generating redundant content regardless of task difficulty. Inspired by the dual process theory in cognitive science, we propose Adaptive Cognition Policy Optimization (ACPO), a reinforcement learning framework that enables LRMs to achieve efficient reasoning through adaptive cognitive allocation and dynamic system switch. ACPO incorporates two key components: (1) introducing system-aware reasoning tokens to explicitly represent the thinking modes thereby making the model's cognitive process transparent, and (2) integrating online difficulty estimation and token length budget to guide adaptive system switch and reasoning during reinforcement learning. To this end, we propose a two-stage training strategy. The first stage begins with supervised fine-tuning to cold start the model, enabling it to generate reasoning paths with explicit thinking modes. In the second stage, we apply ACPO to further enhance adaptive system switch for difficulty-aware reasoning. Experimental results demonstrate that ACPO effectively reduces redundant reasoning while adaptively adjusting cognitive allocation based on task complexity, achieving efficient hybrid reasoning. Recent advances in large reasoning models (LRMs) [1] have demonstrated remarkable success on complex tasks such as mathematical reasoning [2, 3, 4, 5], largely attributed to reinforcement learning that encourages the generation of detailed, step-by-step reasoning processes. LRMs improve answer accuracy through self-reflection and self-verification during long reasoning paths. As the reasoning length increases, the performance of model tends to improve accordingly [2, 6, 7] . Although the long chain-of-thought (CoT) [8] reasoning in LRMs is effective for solving complex problems, it often leads to overthinking [9, 10], producing redundant reasoning paths. Most existing LRMs rely on fixed reasoning strategies, lacking the ability to dynamically switch between different thinking modes based on task complexity. This rigidity results in inefficient inference, particularly for simple problems that could be resolved more effectively with concise and direct reasoning. Several recent efforts have explored long CoT compression for efficient reasoning [11, 12] . One line of work fine-tunes LRMs using supervision from shorter chain-of-thought exemplars [13, 14] , encouraging the model to arrive at correct answers with fewer intermediate steps. Another line introduces length penalties into reinforcement learning reward functions [3, 15, 16, 17] , explicitly * Equal Contribution.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bb1f760f-41b5-4990-b201-75888f64cb3fCited by top-tier papers7
- VideoAuto-R1: Video Auto Reasoning via Thinking Once, Answering TwiceShuming Liu, Mingchen Zhuge, Changsheng Zhao, Jun Chen et al.CVPR 2026 · 18 citations
- DeepCompress: A Dual Reward Strategy for Dynamically Exploring and Compressing Reasoning ChainsTian Liang, Wenxiang Jiao, Zhiwei He, Jiahao Xu et al.ICLR 2026 · 10 citations
- ConPress: Learning Efficient Reasoning from Multi-Question Contextual PressureJie Deng, Shining Liang, Jun Li, Hongzhi Li et al.ICML 2026 · 3 citations
- A State-Transition Framework for Efficient LLM ReasoningLiang Zhang, Yu Zhao, Longyue Wang, Tianqi Shi et al.ICLR 2026 · 2 citations
- ForesightKV: Optimizing KV Cache Eviction for Reasoning Models by Learning Long-Term ContributionZican Dong, Peiyu Liu, Junyi Li, Zhipeng Chen et al.ICML 2026
Builds on7
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- SimPO: Simple Preference Optimization with a Reference-Free RewardYu Meng, Mengzhou Xia, Danqi ChenNeurIPS 2024 · 1,203 citations
- CoT-Valve: Length-Compressible Chain-of-Thought TuningXinyin Ma, Guangnian Wan, Runpeng Yu, Gongfan Fang et al.ACL 2025 · 162 citations
- Towards Thinking-Optimal Scaling of Test-Time Compute for LLM ReasoningWenkai Yang, Shuming Ma, Yankai Lin, Furu WeiNeurIPS 2025 · 141 citations
- Omni-MATH: A Universal Olympiad Level Mathematic Benchmark for Large Language ModelsBofei Gao, Feifan Song, Zhe Yang, Zefan Cai et al.ICLR 2025 · 3 citations
Related papers
- Think Only When You Need with Large Hybrid-Reasoning ModelsLingjie Jiang, Xun Wu, Shaohan Huang, Qingxiu Dong et al.NeurIPS 2025 · 71 citations
- When Simple Problems Wear Complex Costumes: Improving Efficiency in LRM's Adaptive ReasoningJunnan Ren, Yan Zhang, Qian Chen, Yunhang Shen et al.ICML 2026
- SuCo: Sufficiency-guided Continuous Adaptive ReasoningJiahao Wang, Bingyu Liang, Chenhao Hu, Longhui Zhang et al.ICML 2026
- Efficiently Learning To Reason or Not to Reason: Root-token Policy Optimization for Adaptive ThinkingTaehyeon Kim, Hyunsoo Lee, Youngsoo Jang, Moontae LeeACL 2026
- AttnPO: Attention-Guided Process Supervision for Efficient ReasoningShuaiyi Nie, Siyu Ding, Wenyuan Zhang, Linhao Yu et al.ACL 2026 · 21 citations
