SpecExec: Massively Parallel Speculative Decoding For Interactive LLM Inference on Consumer Devices
Ruslan Svirschevski, Avner May, Zhuoming Chen, Beidi Chen, Zhihao Jia, Max Ryabinin
摘要
As large language models gain widespread adoption, running them efficiently becomes crucial. Recent works on LLM inference use speculative decoding to achieve extreme speedups. However, most of these works implicitly design their algorithms for high-end datacenter hardware. In this work, we ask the opposite question: how fast can we run LLMs on consumer machines? Consumer GPUs can no longer fit the largest available models (50B+ parameters) and must offload them to RAM or SSD. When running with offloaded parameters, the inference engine can process batches of hundreds or thousands of tokens at the same time as just one token, making it a natural fit for speculative decoding. We propose SpecExec (Speculative Execution), a simple parallel decoding method that can generate up to 20 tokens per target model iteration for popular LLM families. It utilizes the high spikiness of the token probabilities distribution in modern LLMs and a high degree of alignment between model output probabilities. SpecExec takes the most probable tokens continuation from the draft model to build a"cache"tree for the target model, which then gets validated in a single pass. Using SpecExec, we demonstrate inference of 50B+ parameter LLMs on consumer GPUs with RAM offloading at 4-6 tokens per second with 4-bit quantization or 2-3 tokens per second with 16-bit weights.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time TestYuhui Li, Fangyun Wei, Chao Zhang, Hongyang ZhangNeurIPS 2025 · 被引用 347 次
- MoESD: Unveil Speculative Decoding's Potential for Accelerating Sparse MoEZongle Huang, Lei Zhu, Zongyuan Zhan, Ting Hu 等NeurIPS 2025 · 被引用 23 次
- SpecEdge: Scalable Edge-Assisted Serving Framework for Interactive LLMsJinwoo Park, Seunggeun Cho, Dongsu HanNeurIPS 2025 · 被引用 20 次
- EAGLE-2: Faster Inference of Language Models with Dynamic Draft TreesYuhui Li, Fangyun Wei, Chao Zhang, Hongyang ZhangEMNLP 2024 · 被引用 16 次
- Speculative Speculative DecodingTanishq Kumar, Tri Dao, Avner MayICLR 2026 · 被引用 15 次
它引用的顶会 Paper14
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 被引用 1,472 次
- FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPUYing Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li 等ICML 2023 · 被引用 683 次
- ZeRO-Offload: Democratizing Billion-Scale Model TrainingJie Ren, Samyam Rajbhandari, Reza Yazdani Aminabadi, Olatunji Ruwase 等USENIX ATC 2021 · 被引用 657 次
- QuIP: 2-Bit Quantization of Large Language Models With GuaranteesJerry Chee, Yaohui Cai, Volodymyr Kuleshov, Christopher De SaNeurIPS 2023 · 被引用 503 次
- Sheared LLaMA: Accelerating Language Model Pre-training via Structured PruningMengzhou Xia, Tianyu Gao, Zhiyuan Zeng, Danqi ChenICLR 2024 · 被引用 453 次
相关 Paper
- Speculate Deep and Accurate: Lossless and Training-Free Acceleration for Offloaded LLMs via Substitute Speculative DecodingPei-Shuo Wang, Jian-Jia Chen, Chun-Che Yang, Chi-Chih Chang 等NeurIPS 2025 · 被引用 1 次
- SpecInfer: Accelerating Large Language Model Serving with Tree-based Speculative Inference and VerificationXupeng Miao, Gabriele Oliaro, Zhihao Zhang, Xinhao Cheng 等ASPLOS 2024 · 被引用 105 次
- EasySpec: Layer-Parallel Speculative Decoding for Efficient Multi-GPU UtilizationYize Wu, Ke Gao, Ling Li, Yanjun WuNeurIPS 2025 · 被引用 3 次
- Sequoia: Scalable and Robust Speculative DecodingZhuoming Chen, Avner May, Ruslan Svirschevski, Yuhsun Huang 等NeurIPS 2024 · 被引用 53 次
- SPIN: Accelerating Large Language Model Inference with Heterogeneous Speculative ModelsFahao Chen, Peng Li, Tom H. Luan, Zhou Su 等INFOCOM 2025 · 被引用 10 次
