Draft Model Knows When to Stop: Self-Verification Speculative Decoding for Long-Form Generation
Ziyin Zhang, Jiahao Xu, Tian Liang, Xingyu Chen, Zhiwei He, Rui Wang, Zhaopeng Tu
Abstract
Conventional speculative decoding (SD) methods utilize a predefined length policy for proposing drafts, which implies the premise that the target model smoothly accepts the proposed draft tokens. However, reality deviates from this assumption: the oracle draft length varies significantly, and the fixed-length policy hardly satisfies such a requirement. Moreover, such discrepancy is further exacerbated in scenarios involving complex reasoning and long-form generation, particularly under testtime scaling for reasoning-specialized models. Through both theoretical and empirical estimation, we establish that the discrepancy between the draft and target models can be approximated by the draft model's prediction entropy: a high entropy indicates a low acceptance rate of draft tokens, and vice versa. Based on this insight, we propose SVIP: Self-Verification Length Policy for Long-Context Speculative Decoding, which is a training-free dynamic length policy for speculative decoding systems that adaptively determines the lengths of draft sequences by referring to the draft entropy. Experimental results on mainstream SD benchmarks as well as reasoning-heavy benchmarks demonstrate the superior performance of SVIP, achieving up to 17% speedup on MT-Bench at 8K context compared with fixed draft lengths, and 22% speedup for QwQ in long-form reasoning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- PipeSD: An Efficient Cloud-Edge Collaborative Pipeline Inference Framework with Speculative DecodingHan Yunhe, Yunqi Gao, Bing Hu, Mahdi Boloursaz Mashhadi et al.ICML 2026 · 2 citations
- EDSD: Entropy-Driven Design for Faster Speculative DecodingLongkai Cheng, Ximing Wang, Jiangcai Zhu, Kailai Shao et al.ACL 2026
Builds on9
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 1,472 citations
- Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding HeadsTianle Cai, Yuhong Li, Zhengyang Geng, Hongwu Peng et al.ICML 2024 · 669 citations
- EAGLE: Speculative Sampling Requires Rethinking Feature UncertaintyYuhui Li, Fangyun Wei, Chao Zhang, Hongyang ZhangICML 2024 · 424 citations
- SpecTr: Fast Speculative Decoding via Optimal TransportZiteng Sun, Ananda Theertha Suresh, Jae Hun Ro, Ahmad Beirami et al.NeurIPS 2023 · 164 citations
- GliDe with a CaPE: A Low-Hassle Method to Accelerate Speculative DecodingCunxiao Du, Jing Jiang, Yuanchen Xu, Jiawei Wu et al.ICML 2024 · 72 citations
Related papers
- KnapSpec: Self-Speculative Decoding via Adaptive Layer Selection as a Knapsack ProblemSeongjin Cha, Gyuwan Kim, Dongsu Han, Tao Yang et al.ICML 2026
- LongSpec: Long-Context Lossless Speculative Decoding with Efficient Drafting and VerificationPenghui Yang, Cunxiao Du, Fengzhuo Zhang, Haonan Wang et al.ACL 2026 · 12 citations
- PEARL: Parallel Speculative Decoding with Adaptive Draft LengthTianyu Liu, Yun Li, Qitan Lv, Kai Liu et al.ICLR 2025
- Calibrated Speculative Decoding: Frequency-Guided Candidate Selection for Efficient InferenceXuwen Zhou, Fangxin Liu, Chao Wang, Xiao Zheng et al.ACL 2026
- Training-Free Loosely Speculative Decoding: Accepting Semantically Correct Drafts Beyond Exact MatchJinze Li, Yixing Xu, Guanchen Li, Shuo Yang et al.ICLR 2026 · 12 citations
