Multi-Draft Speculative Sampling: Canonical Decomposition and Theoretical Limits
Ashish J. Khisti, MohammadReza Ebrahimi, Hassan Dbouk, Arash Behboodi, Roland Memisevic, Christos Louizos
Abstract
We consider multi-draft speculative sampling, where the proposal sequences are sampled independently from different draft models. At each step, a token-level draft selection scheme takes a list of valid tokens as input and produces an output token whose distribution matches that of the target model. Previous works have demonstrated that the optimal scheme (which maximizes the probability of accepting one of the input tokens) can be cast as a solution to a linear program. In this work we show that the optimal scheme can be decomposed into a two-step solution: in the first step an importance sampling (IS) type scheme is used to select one intermediate token; in the second step (single-draft) speculative sampling is applied to generate the output token. For the case of two identical draft models we further 1) establish a necessary and sufficient condition on the distributions of the target and draft models for the acceptance probability to equal one and 2) provide an explicit expression for the optimal acceptance probability. Our theoretical analysis also motives a new class of token-level selection schemes based on weighted importance sampling. Our experimental results demonstrate consistent improvements in the achievable block efficiency and token rates over baseline schemes in a number of scenarios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Not-a-Bandit: Provably No-Regret Drafter Selection in Speculative Decoding for LLMsHongyi Liu, Jiaji Huang, Zhen Jia, Youngsuk Park et al.ICLR 2026 · 5 citations
- List-Level Distribution Coupling with Applications to Speculative Decoding and Lossy CompressionJoseph Rowan, Buu Phan, Ashish KhistiNeurIPS 2025 · 4 citations
- Global Resolution: Optimal Multi-Draft Speculative Sampling via Convex OptimizationRahul Krishna Thomas, Arka PalICLR 2026 · 2 citations
Builds on6
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 1,472 citations
- SparseGPT: Massive Language Models Can be Accurately Pruned in One-ShotElias Frantar, Dan AlistarhICML 2023 · 1,240 citations
- Quantizable Transformers: Removing Outliers by Helping Attention Heads Do NothingYelysei Bondarenko, Markus Nagel, Tijmen BlankevoortNeurIPS 2023 · 196 citations
- SpecTr: Fast Speculative Decoding via Optimal TransportZiteng Sun, Ananda Theertha Suresh, Jae Hun Ro, Ahmad Beirami et al.NeurIPS 2023 · 164 citations
Related papers
- Towards Optimal Multi-draft Speculative DecodingZhengmian Hu, Tong Zheng, Vignesh Viswanathan, Ziyi Chen et al.ICLR 2025
- Block Verification Accelerates Speculative DecodingZiteng Sun, Uri Mendlovic, Yaniv Leviathan, Asaf Aharoni et al.ICLR 2025
- Cascade Speculative Drafting for Even Faster LLM InferenceZiyi Chen, Xiaocong Yang, Jiacheng Lin, Chenkai Sun et al.NeurIPS 2024 · 107 citations
- SpecHub: Provable Acceleration to Multi-Draft Speculative DecodingRyan Sun, Tianyi Zhou, Xun Chen, Lichao SunEMNLP 2024
- Gumiho: A Hybrid Architecture to Prioritize Early Tokens in Speculative DecodingJinze Li, Yixing Xu, Haiduo Huang, Xuanwu Yin et al.ICML 2025
