Accelerated Speculative Sampling Based on Tree Monte Carlo
Zhengmian Hu, Heng Huang
Abstract
Speculative Sampling (SpS) has been introduced to speed up the inference of large language models (LLMs) by generating multiple tokens in a single forward pass under the guidance of a reference model, while preserving the original distribution. We observe that SpS can be derived through maximum coupling on the token distribution. However, we find that this approach is not optimal as it applies maximum coupling incrementally for each new token, rather than seeking a global maximum coupling that yields a faster algorithm, given the tree-space nature of LLM generative distributions. In this paper, we shift our focus from distributions on a token space to those on a tree space. We propose a novel class of Tree Monte Carlo (TMC) methods, demonstrating their unbiasedness and convergence. As a particular instance of TMC, our new algorithm, Accelerated Speculative Sampling (ASpS), outperforms traditional SpS by generating more tokens per step on average, achieving faster inference, while maintaining the original distribution. This challenge has prompted the exploration of methods
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a74b9ca3-fd11-4ca1-97cd-a16bc326c1a9Cited by top-tier papers10
- Traversal Verification for Speculative Tree DecodingYepeng Weng, Qiao Hu, Xujie Chen, Li Liu et al.NeurIPS 2025 · 11 citations
- CORAL: Learning Consistent Representations across Multi-step Training with Lighter Speculative DrafterYepeng Weng, Dianwen Mei, Huishi Qiu, Xujie Chen et al.ACL 2025 · 6 citations
- Global Resolution: Optimal Multi-Draft Speculative Sampling via Convex OptimizationRahul Krishna Thomas, Arka PalICLR 2026 · 2 citations
- Annealed Relaxation of Speculative Decoding for Faster Autoregressive Image GenerationXingyao Li, Fengzhuo Zhang, Cunxiao Du, Hui JiAAAI 2026
- Towards Optimal Multi-draft Speculative DecodingZhengmian Hu, Tong Zheng, Vignesh Viswanathan, Ziyi Chen et al.ICLR 2025
Builds on3
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 1,472 citations
- Speculative Decoding with Big Little DecoderSehoon Kim, Karttikeya Mangalam, Suhong Moon, Jitendra Malik et al.NeurIPS 2023 · 212 citations
- SpecInfer: Accelerating Large Language Model Serving with Tree-based Speculative Inference and VerificationXupeng Miao, Gabriele Oliaro, Zhihao Zhang, Xinhao Cheng et al.ASPLOS 2024 · 105 citations
Related papers
- EAGLE-2: Faster Inference of Language Models with Dynamic Draft TreesYuhui Li, Fangyun Wei, Chao Zhang, Hongyang ZhangEMNLP 2024 · 16 citations
- Accelerated Diffusion Models via Speculative SamplingValentin De Bortoli, Alexandre Galashov, Arthur Gretton, Arnaud DoucetICML 2025
- Re-SpS: A Reinforcement Learning Approach to Speculative SamplingChenan Wang, Daniel H. Shi, Haipeng ChenAAAI 2026
- Cactus: Accelerating Auto-Regressive Decoding with Constrained Acceptance Speculative SamplingYongchang Hao, Lili MouICLR 2026 · 3 citations
- Dynamic-Width Speculative Beam Decoding for LLM InferenceZongyue Qin, Zifan He, Neha Prakriya, Jason Cong et al.AAAI 2025 · 10 citations
