Accelerated Speculative Sampling Based on Tree Monte Carlo
Zhengmian Hu, Heng Huang
摘要
Speculative Sampling (SpS) has been introduced to speed up the inference of large language models (LLMs) by generating multiple tokens in a single forward pass under the guidance of a reference model, while preserving the original distribution. We observe that SpS can be derived through maximum coupling on the token distribution. However, we find that this approach is not optimal as it applies maximum coupling incrementally for each new token, rather than seeking a global maximum coupling that yields a faster algorithm, given the tree-space nature of LLM generative distributions. In this paper, we shift our focus from distributions on a token space to those on a tree space. We propose a novel class of Tree Monte Carlo (TMC) methods, demonstrating their unbiasedness and convergence. As a particular instance of TMC, our new algorithm, Accelerated Speculative Sampling (ASpS), outperforms traditional SpS by generating more tokens per step on average, achieving faster inference, while maintaining the original distribution. This challenge has prompted the exploration of methods
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Traversal Verification for Speculative Tree DecodingYepeng Weng, Qiao Hu, Xujie Chen, Li Liu 等NeurIPS 2025 · 被引用 11 次
- CORAL: Learning Consistent Representations across Multi-step Training with Lighter Speculative DrafterYepeng Weng, Dianwen Mei, Huishi Qiu, Xujie Chen 等ACL 2025 · 被引用 6 次
- Global Resolution: Optimal Multi-Draft Speculative Sampling via Convex OptimizationRahul Krishna Thomas, Arka PalICLR 2026 · 被引用 2 次
- Annealed Relaxation of Speculative Decoding for Faster Autoregressive Image GenerationXingyao Li, Fengzhuo Zhang, Cunxiao Du, Hui JiAAAI 2026
- Towards Optimal Multi-draft Speculative DecodingZhengmian Hu, Tong Zheng, Vignesh Viswanathan, Ziyi Chen 等ICLR 2025
它引用的顶会 Paper3
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 被引用 1,472 次
- Speculative Decoding with Big Little DecoderSehoon Kim, Karttikeya Mangalam, Suhong Moon, Jitendra Malik 等NeurIPS 2023 · 被引用 212 次
- SpecInfer: Accelerating Large Language Model Serving with Tree-based Speculative Inference and VerificationXupeng Miao, Gabriele Oliaro, Zhihao Zhang, Xinhao Cheng 等ASPLOS 2024 · 被引用 105 次
相关 Paper
- EAGLE-2: Faster Inference of Language Models with Dynamic Draft TreesYuhui Li, Fangyun Wei, Chao Zhang, Hongyang ZhangEMNLP 2024 · 被引用 16 次
- Accelerated Diffusion Models via Speculative SamplingValentin De Bortoli, Alexandre Galashov, Arthur Gretton, Arnaud DoucetICML 2025
- Re-SpS: A Reinforcement Learning Approach to Speculative SamplingChenan Wang, Daniel H. Shi, Haipeng ChenAAAI 2026
- Cactus: Accelerating Auto-Regressive Decoding with Constrained Acceptance Speculative SamplingYongchang Hao, Lili MouICLR 2026 · 被引用 3 次
- Dynamic-Width Speculative Beam Decoding for LLM InferenceZongyue Qin, Zifan He, Neha Prakriya, Jason Cong 等AAAI 2025 · 被引用 10 次
