VeriThinker: Learning to Verify Makes Reasoning Model Efficient
Zigeng Chen, Xinyin Ma, Gongfan Fang, Ruonan Yu, Xinchao Wang
Abstract
Large Reasoning Models (LRMs) excel at complex tasks using Chain-of-Thought (CoT) reasoning. However, their tendency to overthinking leads to unnecessarily lengthy reasoning chains, dramatically increasing inference costs. To mitigate this issue, we introduce VeriThinker, a novel approach for CoT compression. Unlike conventional methods that fine-tune LRMs directly on the original reasoning task using synthetic concise CoT data, we innovatively fine-tune the model solely through an auxiliary verification task. By training LRMs to accurately verify the correctness of CoT solutions, the LRMs inherently become more discerning about the necessity of subsequent self-reflection steps, thereby effectively suppressing overthinking. Extensive experiments validate that VeriThinker substantially reduces reasoning chain lengths while maintaining or even slightly improving accuracy. When applied to DeepSeek-R1-Distill-Qwen-7B, our approach reduces reasoning tokens on MATH500 from 3790 to 2125 while improving accuracy by 0.8% (94.0% to 94.8%), and on AIME25, tokens decrease from 14321 to 10287 with a 2.1% accuracy gain (38.7% to 40.8%). Additionally, our experiments demonstrate that VeriThinker can also be zero-shot generalized to speculative reasoning. Code is available at https://github.com/czg1225/VeriThinker
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8a5fcb8b-abd8-436d-abab-1684a47a9643Cited by top-tier papers5
- OptimalThinkingBench: Evaluating Over and Underthinking in LLMsPranjal Aggarwal, Seungone Kim, Jack Lanchantin, Sean Welleck et al.ICLR 2026 · 38 citations
- Render-of-Thought: Rendering Textual Chain-of-Thought as Images for Visual Latent ReasoningYifan Wang, Shiyu Li, Peiming Li, Xiaochen Yang et al.ACL 2026 · 14 citations
- Learning to Self-Verify Makes Language Models Better ReasonersYuxin Chen, Yu Wang, Yi Zhang, Ziang Ye et al.ICML 2026 · 12 citations
- Your Reasoning Model Knows What Counts: Self-Guided Chain-of-Thought Pruning for Efficient ReasoningZi-Ao Ma, Xian-Ling Mao, Tian Lan, Chen Xu et al.ACL 2026
- Stop When Further Reasoning Won’t Help: Attention-State Adaptive Generation in Reasoning ModelsJiakai Li, KE QIN, Rongzheng Wang, Yizhuo Ma et al.ICML 2026
Builds on32
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 1,472 citations
Related papers
- ConPress: Learning Efficient Reasoning from Multi-Question Contextual PressureJie Deng, Shining Liang, Jun Li, Hongzhi Li et al.ICML 2026 · 3 citations
- ThoughtFold: Folding Reasoning Chains via Introspective Preference LearningZiyan Liu, Xueda Shen, Yuzhe Gu, Songyang Gao et al.ICML 2026 · 3 citations
- TrimR: Verifier-based Training-Free Thinking Trimming for Efficient Test-Time ScalingWeizhe Lin, Xing Li 023, Zhiyuan Yang, Xiaojin Fu et al.ICLR 2026 · 14 citations
- Think Better, Not Longer: Token-Level Marginal Utility for Efficient Reasoning in Large Reasoning ModelsJiawei Li, Yang Gao, Huashan Sun, Chong FengACL 2026
- TokenSqueeze: Performance-Preserving Compression for Reasoning LLMsYuxiang Zhang, Zhengxu Yu, Weihang Pan, Zhongming Jin et al.NeurIPS 2025 · 6 citations
