TrimR: Verifier-based Training-Free Thinking Trimming for Efficient Test-Time Scaling
Weizhe Lin, Xing Li 023, Zhiyuan Yang, Xiaojin Fu, Hui-Ling Zhen, Yaoyuan Wang, Xianzhi Yu, Wulong Liu, Xiaosong Li, Mingxuan Yuan
Abstract
Large Reasoning Models (LRMs) demonstrate exceptional capability in tackling complex mathematical, logical, and coding tasks by leveraging extended Chain-of-Thought (CoT) reasoning. Test-time scaling methods-such as prolonging CoT with explicit token-level exploration-can push LRMs' accuracy boundaries, but they incur significant decoding overhead. A key inefficiency source is LRMs often generate redundant thinking CoTs, which demonstrate clear structured overthinking and underthinking patterns. Inspired by human cognitive reasoning processes and numerical optimization theories, we propose TrimR, a verifier-based, trainingfree, efficient framework for dynamic CoT compression to trim reasoning and enhance test-time scaling, explicitly tailored for production-level deployment. Our method employs a lightweight, pretrained, instruction-tuned verifier to detect and truncate redundant intermediate thoughts of LRMs without any LRM or verifier fine-tuning. We present both the core algorithm and asynchronous online system engineered for high-throughput industrial applications. Empirical evaluations on Ascend NPUs and vLLM show that our framework delivers substantial gains in inference efficiency under large-batch workloads. In particular, on the four MATH500, AIME24/25, and GPQA benchmarks, the reasoning runtime of Pangu Pro MoE, Pangu-R-38B, QwQ-32B, and DeepSeek-R1-Distill-Qwen-32B is improved by up to 70% with negligible impact on accuracy. <think>, so I have this problem about Aya's morning walk ... Let me try to parse it step by step. ... , ... so total time is 3 hours 24 minutes, which is 204 minutes. So is that the answer? , ... so 204 is correct. Hmm. Alternatively, maybe I made a mistake in interpreting the total time? , maybe I need to present the answer in hours converted to minutes ..., so 204 is the answer.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 36ff0bb6-4221-44c7-9b99-3511ecd15e4fCited by top-tier papers2
- Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative ExplorationShuzhang Zhong, Haochen Huang, Shengxuan Qiu, Pengfei Zuo et al.OSDI 2026
- Learning Graph Rationales to Compress Long Chains of Thought in Multimodal ReasoningYizhi Wang, Linan Yue, Deng-Bao Wang, Tong Wei et al.KDD 2026
Builds on11
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- SGLang: Efficient Execution of Structured Language Model ProgramsLianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun et al.NeurIPS 2024 · 1,586 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- Training Language Models to Reason EfficientlyDaman Arora, Andrea ZanetteNeurIPS 2025 · 270 citations
- Dynamic Early Exit in Reasoning ModelsChenxu Yang, Qingyi Si, Yongjie Duan, Zheliang Zhu et al.ICLR 2026 · 250 citations
Related papers
- VeriThinker: Learning to Verify Makes Reasoning Model EfficientZigeng Chen, Xinyin Ma, Gongfan Fang, Ruonan Yu et al.NeurIPS 2025 · 30 citations
- Adaptive Spatial and Temporal Redundancy Optimization for Efficient Reasoning in Large Language ModelsTianle Chen, Pengyu Cheng, Qiyuan Zhu, Jiacheng Wang et al.ACL 2026 · 1 citation
- Do NOT Think That Much for 2+3=? On the Overthinking of Long Reasoning ModelsXingyu Chen, Jiahao Xu, Tian Liang, Zhiwei He et al.ICML 2025
- LIMOPro: Reasoning Refinement for Efficient and Effective Test-time ScalingYang Xiao, Jiashuo Wang, Ruifeng Yuan, Chunpu Xu et al.NeurIPS 2025 · 15 citations
- SLAT: Segment-Level Adaptive Trimming for Efficient CoT ReasoningJian Yao, Xiongcai Luo, Ran Cheng, KC TanICML 2026
