Towards Optimal Multi-draft Speculative Decoding
Zhengmian Hu, Tong Zheng, Vignesh Viswanathan, Ziyi Chen, Ryan A. Rossi, Yihan Wu, Dinesh Manocha, Heng Huang
Abstract
Large Language Models (LLMs) have become an indispensable part of natural language processing tasks. However, autoregressive sampling has become an efficiency bottleneck. Multi-Draft Speculative Decoding (MDSD) is a recent approach where, when generating each token, a small draft model generates multiple drafts, and the target LLM verifies them in parallel, ensuring that the final output conforms to the target model distribution. The two main design choices in MDSD are the draft sampling method and the verification algorithm. For a fixed draft sampling method, the optimal acceptance rate is a solution to an optimal transport problem, but the complexity of this problem makes it difficult to solve for the optimal acceptance rate and measure the gap between existing verification algorithms and the theoretical upper bound. This paper discusses the dual of the optimal transport problem, providing a way to efficiently compute the optimal acceptance rate. For the first time, we measure the theoretical upper bound of MDSD efficiency for vocabulary sizes in the thousands and quantify the gap between existing verification algorithms and this bound. We also compare different draft sampling methods based on their optimal acceptance rates. Our results show that the draft sampling method strongly influences the optimal acceptance rate, with sampling without replacement outperforming sampling with replacement. Additionally, existing verification algorithms do not reach the theoretical upper bound for both without replacement and with replacement sampling. Our findings suggest that carefully designed draft sampling methods can potentially improve the optimal acceptance rate and enable the development of verification algorithms that closely match the theoretical upper bound.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- Traversal Verification for Speculative Tree DecodingYepeng Weng, Qiao Hu, Xujie Chen, Li Liu et al.NeurIPS 2025 · 11 citations
- TETRIS: Optimal Draft Token Selection for Batch Speculative DecodingZhaoxuan Wu, Zijian Zhou, Arun Verma, Alok Prakash et al.ACL 2025 · 7 citations
- Overcoming Joint Intractability with Lossless Hierarchical Speculative DecodingYuxuan Zhou, Fei Huang, Heng Li, Fengyi Wu et al.ICLR 2026 · 5 citations
- Not-a-Bandit: Provably No-Regret Drafter Selection in Speculative Decoding for LLMsHongyi Liu, Jiaji Huang, Zhen Jia, Youngsuk Park et al.ICLR 2026 · 5 citations
- SelfJudge: Faster Speculative Decoding via Self-Supervised Judge VerificationKanghoon Yoon, Minsub Kim, Sungjae Lee, Joonhyung Lee et al.ICML 2026 · 4 citations
Builds on7
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 1,472 citations
- SpecTr: Fast Speculative Decoding via Optimal TransportZiteng Sun, Ananda Theertha Suresh, Jae Hun Ro, Ahmad Beirami et al.NeurIPS 2023 · 164 citations
- DistillSpec: Improving Speculative Decoding via Knowledge DistillationYongchao Zhou, Kaifeng Lyu, Ankit Singh Rawat, Aditya Krishna Menon et al.ICLR 2024 · 143 citations
- Cascade Speculative Drafting for Even Faster LLM InferenceZiyi Chen, Xiaocong Yang, Jiacheng Lin, Chenkai Sun et al.NeurIPS 2024 · 107 citations
Related papers
- Global Resolution: Optimal Multi-Draft Speculative Sampling via Convex OptimizationRahul Krishna Thomas, Arka PalICLR 2026 · 2 citations
- SpecHub: Provable Acceleration to Multi-Draft Speculative DecodingRyan Sun, Tianyi Zhou, Xun Chen, Lichao SunEMNLP 2024
- Accelerating LLM Inference with Lossless Speculative Decoding Algorithms for Heterogeneous VocabulariesNadav Timor, Jonathan Mamou, Daniel Korat, Moshe Berchansky et al.ICML 2025
- A Theoretical Perspective for Speculative Decoding AlgorithmMing Yin, Minshuo Chen, Kaixuan Huang, Mengdi WangNeurIPS 2024 · 36 citations
- Dynamic-Width Speculative Beam Decoding for LLM InferenceZongyue Qin, Zifan He, Neha Prakriya, Jason Cong et al.AAAI 2025 · 10 citations
