polybasic Speculative Decoding Through a Theoretical Perspective
Ruilin Wang, Huixia Li, Yuexiao Ma, Xiawu Zheng, Fei Chao, Xuefeng Xiao, Rongrong Ji
Abstract
Inference latency stands as a critical bottleneck in the large-scale deployment of Large Language Models (LLMs). Speculative decoding methods have recently shown promise in accelerating inference without compromising the output distribution. However, existing work typically relies on a dualistic draft-verify framework and lacks rigorous theoretical grounding. In this paper, we introduce a novel polybasic speculative decoding framework, underpinned by a comprehensive theoretical analysis. Specifically, we prove a fundamental theorem that characterizes the optimal inference time for multi-model speculative decoding systems, shedding light on how to extend beyond the dualistic approach to a more general polybasic paradigm. Through our theoretical investigation of multi-model token generation, we expose and optimize the interplay between model capabilities, acceptance lengths, and overall computational cost. Our framework supports both standalone implementation and integration with existing speculative techniques, leading to accelerated performance in practice. Experimental results across multiple model families demonstrate that our approach yields speedup ratios ranging from 3.31× to 4.01× for LLaMA2-Chat 7B, up to 3.87× for LLaMA3-8B, up to 4.43× for Vicuna-7B and up to 3.85× for Qwen2-7B-all while preserving the original output distribution. We release our theoretical proofs and implementation code to facilitate further investigation into polybasic speculative decoding.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4f21aacf-a1d4-4e2f-8d41-9a21999a1928Builds on16
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 1,472 citations
- H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language ModelsZhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen et al.NeurIPS 2023 · 1,003 citations
- Speculative Decoding with Big Little DecoderSehoon Kim, Karttikeya Mangalam, Suhong Moon, Jitendra Malik et al.NeurIPS 2023 · 212 citations
- LLM Maybe LongLM: SelfExtend LLM Context Window Without TuningHongye Jin, Xiaotian Han, Jingfeng Yang, Zhimeng Jiang et al.ICML 2024 · 167 citations
Related papers
- Falcon: Faster and Parallel Inference of Large Language Models Through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding TreeXiangxiang Gao, Weisheng Xie, Yiwei Xiang, Feng JiAAAI 2025 · 20 citations
- Talon: Breaking the Synchronization Barrier in Speculative Decoding with Hybrid Model-based and Retrieve-based DraftingXiangxiang Gao, Weisheng Xie, Lixin, Xuwei Fang et al.AAAI 2026
- GliDe with a CaPE: A Low-Hassle Method to Accelerate Speculative DecodingCunxiao Du, Jing Jiang, Yuanchen Xu, Jiawei Wu et al.ICML 2024 · 72 citations
- Sequoia: Scalable and Robust Speculative DecodingZhuoming Chen, Avner May, Ruslan Svirschevski, Yuhsun Huang et al.NeurIPS 2024 · 53 citations
- GRIFFIN: Effective Token Alignment for Faster Speculative DecodingShijing Hu, Jingyang Li, Xingyu Xie, Zhihui Lu et al.NeurIPS 2025 · 14 citations
