From "Aha Moments" to Controllable Thinking: Toward Meta-Cognitive Reasoning in LRMs via Decoupled Reasoning and Control
Rui Ha, Rui Pu, Chaozhuo Li, Li Sun, Sen Su
Abstract
Large Reasoning Models (LRMs) can exhibit step-by-step reasoning, reflection, and backtracking, but these behaviors are often unregulated, leading to overthinking. As a result, LRMs continue generating redundant reasoning even after reaching high-confidence conclusions. This increases inference cost and latency, limiting practical deployment. The root cause is the absence of an intrinsic mechanism to monitor the reasoning state and decide when to continue, backtrack, or stop. We propose MERA, a meta-cognitive reasoning framework that decouples reasoning from control to enable independent optimization of control strategies. MERA constructs high-quality reasoning-control supervision data via a takeover-based pipeline, and transforms long-horizon traces into structured reasoning-control alternating sequences for training. The model is trained with supervised fine-tuning to internalize the structured separation, and further optimized with Control-Segment Policy Optimization (CSPO), which combines segment-wise GRPO with control masking to focus learning on control segments. Experiments across reasoning benchmarks show that MERA improves both efficiency and accuracy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7e8df051-a9a4-4430-b23a-d7d936220b92Builds on8
- Dynamic Early Exit in Reasoning ModelsChenxu Yang, Qingyi Si, Yongjie Duan, Zheliang Zhu et al.ICLR 2026 · 250 citations
- Thinkless: LLM Learns When to ThinkGongfan Fang, Xinyin Ma, Xinchao WangNeurIPS 2025 · 128 citations
- Think Only When You Need with Large Hybrid-Reasoning ModelsLingjie Jiang, Xun Wu, Shaohan Huang, Qingxiu Dong et al.NeurIPS 2025 · 71 citations
- Demystifying Reasoning Dynamics with Mutual Information: Thinking Tokens are Information Peaks in LLM ReasoningChen Qian, Dongrui Liu, Haochen Wen, Zhen Bai et al.NeurIPS 2025 · 63 citations
- Learning on Large-scale Text-attributed Graphs via Variational InferenceJianan Zhao, Meng Qu, Chaozhuo Li, Hao Yan et al.ICLR 2023 · 25 citations
Related papers
- Intervene When It Doubts: Conjunction-Guided Interactive ReasoningQianyue Wang, Jinwu Hu, Yaofo Chen, Yufeng Wang et al.ICML 2026
- Incentivizing Dual Process Thinking for Efficient Large Language Model ReasoningXiaoxue Cheng, Junyi Li, Zhenduo Zhang, Xinyu Tang et al.NeurIPS 2025 · 25 citations
- SuCo: Sufficiency-guided Continuous Adaptive ReasoningJiahao Wang, Bingyu Liang, Chenhao Hu, Longhui Zhang et al.ICML 2026
- Efficiently Learning To Reason or Not to Reason: Root-token Policy Optimization for Adaptive ThinkingTaehyeon Kim, Hyunsoo Lee, Youngsoo Jang, Moontae LeeACL 2026
- When Simple Problems Wear Complex Costumes: Improving Efficiency in LRM's Adaptive ReasoningJunnan Ren, Yan Zhang, Qian Chen, Yunhang Shen et al.ICML 2026
