Trade-offs in Large Reasoning Models: An Empirical Analysis of Deliberative and Adaptive Reasoning over Foundational Capabilities
Weixiang Zhao, Xingyu Sui, Jiahe Guo, Yulin Hu, Yang Deng, Yanyan Zhao, Xuda Zhi, Yongbo Huang, Hao He, Wanxiang Che, Ting Liu, Bing Qin
Abstract
Recent advancements in Large Reasoning Models (LRMs), such as OpenAI's o1/o3 and DeepSeek-R1, have demonstrated remarkable performance in specialized reasoning tasks through human-like deliberative thinking and long chain-ofthought reasoning. However, our systematic evaluation across various model families (DeepSeek, Qwen, and LLaMA) and scales (7B to 32B) reveals that acquiring these deliberative reasoning capabilities significantly reduces the foundational capabilities of LRMs, including notable declines in helpfulness and harmlessness, alongside substantially increased inference costs. Importantly, we demonstrate that adaptive reasoningemploying modes like Zero-Thinking, Less-Thinking, and Summary-Thinking-can effectively alleviate these drawbacks. Our empirical insights underline the critical need for developing more versatile LRMs capable of dynamically allocating inference-time compute according to specific task characteristics. Our code is available at: https://github.com/SCIR- SC-Qiaoban-Team/FreeEvalLM. WARNING: This paper may contain content that is offensive and harmful.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- When Less Language is More: Language-Reasoning Disentanglement Makes LLMs Better Multilingual ReasonersWeixiang Zhao, Jiahe Guo, Yang Deng, Tongtong Wu et al.NeurIPS 2025 · 20 citations
- On Reasoning Strength Planning in Large Reasoning ModelsLeheng Sheng, An Zhang, Zijian Wu, Weixiang Zhao et al.NeurIPS 2025 · 17 citations
- Thinking in Character: Advancing Role-Playing Agents with Role-Aware ReasoningYihong Tang, Kehai Chen, Muyun Yang, Zheng-Yu Niu et al.NeurIPS 2025 · 16 citations
- BARREL: Boundary-Aware Reasoning for Factual and Reliable LRMsJunxiao Yang, Jinzhe Tu, Haoran Liu, Xiaoce Wang et al.ICLR 2026 · 9 citations
- FireScope: Wildfire Risk Raster Prediction With a Chain-of-Thought OracleMario Markov, Stefan Maria Ailuro, Luc Van Gool, Konrad Schindler et al.CVPR 2026
Builds on5
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language ModelsLiwei Jiang, Kavel Rao, Seungju Han, Allyson Ettinger et al.NeurIPS 2024 · 247 citations
- The First Few Tokens Are All You Need: An Efficient and Effective Unsupervised Prefix Fine-Tuning Method for Reasoning ModelsKe Ji, Jiahao Xu, Tian Liang, Qiuzhi Liu et al.NeurIPS 2025 · 33 citations
- SAPT: A Shared Attention Framework for Parameter-Efficient Continual Learning of Large Language ModelsWeixiang Zhao, Shilong Wang, Yulin Hu, Yanyan Zhao et al.ACL 2024
Related papers
- Training Language Models to Reason EfficientlyDaman Arora, Andrea ZanetteNeurIPS 2025 · 270 citations
- Finding and Reactivating Post-Trained LLMs' Hidden Safety MechanismsMingjie Li, Wai Man Si, Michael Backes, Yang Zhang et al.NeurIPS 2025 · 4 citations
- AdaptThink: Reasoning Models Can Learn When to ThinkJiajie Zhang, Nianyi Lin, Lei Hou, Ling Feng et al.EMNLP 2025 · 3 citations
- Thinker: Learning to Think Fast and SlowStephen Chung, Wenyu Du, Jie FuNeurIPS 2025 · 10 citations
- Thoughts Are All Over the Place: On the Underthinking of Long Reasoning ModelsYue Wang, Qiuzhi Liu, Jiahao Xu, Tian Liang et al.NeurIPS 2025 · 13 citations
