SpecExit: Accelerating Large Reasoning Model via Speculative Exit
Rubing Yang, Huajun Bai, Song Liu, Guanghua Yu, Runzhi Fan, Yanbin Dang, Zhang Jiejing, Kai Liu, Jianchen Zhu, Peng Chen
Abstract
Despite their strong performance on reasoning tasks, Large reasoning models (LRMs) often suffer from overthinking, producing unnecessarily long outputs and incurring high end-to-end latency, a significant limitation to their real-world deployment. To address overthinking, early-exit mechanisms have been proposed to terminate reasoning before typical completion, showing that this approach can effectively shorten generation length with minimal impact on accuracy. However, their reliance on probing mechanisms introduces a detection overhead that limits their end-to-end latency gains and compromises their generalizability across diverse problems. Inspired by the use of hidden states in speculative decoding, we propose SpecExit, a novel framework that predicts both future tokens and an early-exit signal directly from a lightweight draft model without probing overhead. Our method offers significant improvements, achieving up to 66% generation length reduction and 2.5 end-to-end speedup compared with the speculative decoding baseline, without compromising accuracy. Our method leverages the inherent signals from hidden states to provide effective early-exit signals, suggesting broader use of hidden states for efficient reasoning. Our code is available at: https://anonymous.4open.science/r/SpecExit-B802.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f696ec80-b30e-470a-b8b7-6927c3583b28Builds on14
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 2,317 citations
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 1,472 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- EAGLE: Speculative Sampling Requires Rethinking Feature UncertaintyYuhui Li, Fangyun Wei, Chao Zhang, Hongyang ZhangICML 2024 · 424 citations
Related papers
- SpecReason: Fast and Accurate Inference-Time Compute via Speculative ReasoningRui Pan, Yinwei Dai, Zhihao Zhang, Gabriele Oliaro et al.NeurIPS 2025 · 68 citations
- Think Faster Than Words: Efficient LLM Chain-of-Thought Reasoning via Dynamic Shortcut DecodingFan Liu, Yanhao Wang, Min Zhang, Zhikang Chen et al.ACL 2026
- Taming the Long-Tail: Efficient Reasoning RL Training with Adaptive DrafterQinghao Hu, Shang Yang, Junxian Guo, Xiaozhe Yao et al.ASPLOS 2026 · 1 citation
- Accelerating Large Language Model Reasoning via Speculative SearchZhihai Wang, Jie Wang, Jilai Pan, Xilin Xia et al.ICML 2025
- When Is Thinking Enough? Early Exit via Sufficiency Assessment for Efficient ReasoningYang Xiang, Yixin Ji, Ruotao Xu, Dan Qiao et al.ACL 2026 · 3 citations
