SpecHub: Provable Acceleration to Multi-Draft Speculative Decoding
Ryan Sun, Tianyi Zhou, Xun Chen, Lichao Sun
摘要
Large Language Models (LLMs) have become essential in advancing natural language processing (NLP) tasks, but their sequential token generation limits inference speed. Multi-Draft Speculative Decoding (MDSD) offers a promising solution by using a smaller draft model to generate multiple token sequences, which the target LLM verifies in parallel. However, current heuristic approaches, such as Recursive Rejection Sampling (RRS), suffer from low acceptance rates in subsequent drafts, limiting the advantages of using multiple drafts. Meanwhile, Optimal Transport with Membership Cost (OTM) can theoretically improve acceptance rates, but its computational cost is too high for real-time use. We present SpecHub, a novel, efficient samplingverification method for MDSD that improves acceptance rates with only linear computational overhead. By simplifying the OTM problem into a compact Linear Programming model, SpecHub significantly reduces computational complexity. It further accelerates sampling by leveraging a sparse joint distribution, focusing computation on high-probability token sequences. In extensive experiments, Spechub consistently generates 0.05-0.27 and 0.02-0.16 more tokens per step than RRS and RRS without replacement. We attach our code at https://github.com/MasterGodzilla/ Speculative_decoding_OT .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Training-Free Loosely Speculative Decoding: Accepting Semantically Correct Drafts Beyond Exact MatchJinze Li, Yixing Xu, Guanchen Li, Shuo Yang 等ICLR 2026 · 被引用 12 次
- Traversal Verification for Speculative Tree DecodingYepeng Weng, Qiao Hu, Xujie Chen, Li Liu 等NeurIPS 2025 · 被引用 11 次
- CopySpec: Accelerating LLMs with Speculative Copy-and-PasteRazvan-Gabriel Dumitru, Minglai Yang, Vikas Yadav, Mihai SurdeanuEMNLP 2025
- Towards Optimal Multi-draft Speculative DecodingZhengmian Hu, Tong Zheng, Vignesh Viswanathan, Ziyi Chen 等ICLR 2025
它引用的顶会 Paper9
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 被引用 1,472 次
- Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding HeadsTianle Cai, Yuhong Li, Zhengyang Geng, Hongwu Peng 等ICML 2024 · 被引用 669 次
- EAGLE: Speculative Sampling Requires Rethinking Feature UncertaintyYuhui Li, Fangyun Wei, Chao Zhang, Hongyang ZhangICML 2024 · 被引用 424 次
- Break the Sequential Dependency of LLM Inference Using Lookahead DecodingYichao Fu, Peter Bailis, Ion Stoica, Hao ZhangICML 2024 · 被引用 290 次
- SpecTr: Fast Speculative Decoding via Optimal TransportZiteng Sun, Ananda Theertha Suresh, Jae Hun Ro, Ahmad Beirami 等NeurIPS 2023 · 被引用 164 次
相关 Paper
- Global Resolution: Optimal Multi-Draft Speculative Sampling via Convex OptimizationRahul Krishna Thomas, Arka PalICLR 2026 · 被引用 2 次
- Towards Efficient LLM Inference via Collective and Adaptive Speculative DecodingSiqi Wang, Hailong Yang, Xuezhu Wang, Tongxuan Liu 等SC 2025 · 被引用 3 次
- Dynamic-Width Speculative Beam Decoding for LLM InferenceZongyue Qin, Zifan He, Neha Prakriya, Jason Cong 等AAAI 2025 · 被引用 10 次
- SpecBranch: Speculative Decoding via Hybrid Drafting and Rollback-Aware Branch ParallelismYuhao Shen, Junyi Shen, Quan Kong, Tianyu Liu 等ICLR 2026 · 被引用 16 次
- A Theoretical Perspective for Speculative Decoding AlgorithmMing Yin, Minshuo Chen, Kaixuan Huang, Mengdi WangNeurIPS 2024 · 被引用 36 次
