ARM: Adaptive Reasoning Model
Siye Wu, Jian Xie, Yikai Zhang, Aili Chen, Kai Zhang, Yu Su, Yanghua Xiao
Abstract
While large reasoning models demonstrate strong performance on complex tasks, they lack the ability to adjust reasoning token usage based on task difficulty. This often leads to the"overthinking"problem -- excessive and unnecessary reasoning -- which, although potentially mitigated by human intervention to control the token budget, still fundamentally contradicts the goal of achieving fully autonomous AI. In this work, we propose Adaptive Reasoning Model (ARM), a reasoning model capable of adaptively selecting appropriate reasoning formats based on the task at hand. These formats include three efficient ones -- Direct Answer, Short CoT, and Code -- as well as a more elaborate format, Long CoT. To train ARM, we introduce Ada-GRPO, an adaptation of Group Relative Policy Optimization (GRPO), which addresses the format collapse issue in traditional GRPO. Ada-GRPO enables ARM to achieve high token efficiency, reducing tokens by an average of 30%, and up to 70%, while maintaining performance comparable to the model that relies solely on Long CoT. Furthermore, not only does it improve inference efficiency through reduced token generation, but it also brings a 2x speedup in training. In addition to the default Adaptive Mode, ARM supports two additional reasoning modes: 1) Instruction-Guided Mode, which allows users to explicitly specify the reasoning format via special tokens -- ideal when the appropriate format is known for a batch of tasks. 2) Consensus-Guided Mode, which aggregates the outputs of the three efficient formats and resorts to Long CoT in case of disagreement, prioritizing performance with higher token usage.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a89e11f6-8b03-4687-b71e-321b16c76f81Cited by top-tier papers10
- Geometric-Mean Policy OptimizationYuzhong Zhao, Yue Liu, Junpeng Liu, Jingye Chen et al.ICLR 2026 · 104 citations
- Counterfactual VLA: Self-Reflective Vision-Language-Action Model with Adaptive ReasoningZhenghao Peng, Wenhao Ding, Yurong You, Yuxiao Chen et al.CVPR 2026 · 25 citations
- Promoting Efficient Reasoning with Verifiable Stepwise RewardChuhuai Yue, Chengqi Dong, Yinan Gao, Hang He et al.AAAI 2026 · 19 citations
- CodeDance: A Dynamic Tool-integrated MLLM for Executable Visual ReasoningQi Song, Honglin Li, Yingchen Yu, Haoyi Zhou et al.CVPR 2026 · 16 citations
- NeuReasoner: Towards Explainable, Controllable, and Unified Reasoning via Mixture-of-NeuronsHaonan Dong, Kehan Jiang, Haoran Ye, Wenhao Zhu et al.ACL 2026 · 15 citations
Builds on12
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base ModelJingcheng Hu, Yinmin Zhang, Qi Han, Daxin Jiang et al.NeurIPS 2025 · 533 citations
- Training Language Models to Reason EfficientlyDaman Arora, Andrea ZanetteNeurIPS 2025 · 270 citations
Related papers
- Incentivizing Dual Process Thinking for Efficient Large Language Model ReasoningXiaoxue Cheng, Junyi Li, Zhenduo Zhang, Xinyu Tang et al.NeurIPS 2025 · 25 citations
- When Simple Problems Wear Complex Costumes: Improving Efficiency in LRM's Adaptive ReasoningJunnan Ren, Yan Zhang, Qian Chen, Yunhang Shen et al.ICML 2026
- SuCo: Sufficiency-guided Continuous Adaptive ReasoningJiahao Wang, Bingyu Liang, Chenhao Hu, Longhui Zhang et al.ICML 2026
- MARS: Multimodal Adaptive Reasoning Model for Avoiding OverthinkingTan Yue, Qiong Wu, Dongyan ZhaoAAAI 2026
- Thinkless: LLM Learns When to ThinkGongfan Fang, Xinyin Ma, Xinchao WangNeurIPS 2025 · 128 citations
