Faster Cascades via Speculative Decoding
Harikrishna Narasimhan, Wittawat Jitkrittum, Ankit Singh Rawat, Seungyeon Kim, Neha Gupta, Aditya Krishna Menon, Sanjiv Kumar
Abstract
Cascades and speculative decoding are two common approaches to improving language models' inference efficiency. Both approaches interleave two models of different sizes, but via fundamentally distinct mechanisms: cascades employ a deferral rule that invokes the larger model only for "hard" inputs, while speculative decoding uses speculative execution to primarily invoke the larger model in parallel scoring mode. These mechanisms offer different benefits: cascades offer compelling cost-quality trade-offs, often even outperforming the large model; speculative cascades offer impressive speed-ups, while guaranteeing quality-neutrality. In this paper, we leverage the best of both these approaches by designing new speculative cascading techniques that implement their deferral rule through speculative execution. We characterize the optimal deferral rule for our speculative cascades, and employ a plug-in approximation to the optimal rule. Experiments with Gemma and T5 models on a range of language benchmarks show that our approach yields better cost-quality trade-offs than cascading and speculative decoding baselines. INTRODUCTION Large language models (LLMs) have yielded significant advances in quality on a range of natural language processing tasks (
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3f8a8788-0d93-4527-bc26-058a44fe545dCited by top-tier papers21
- Accelerating Diffusion LLMs via Adaptive Parallel DecodingDaniel Israel, Guy Van den Broeck, Aditya GroverNeurIPS 2025 · 114 citations
- Universal Model Routing for Efficient LLM InferenceWittawat Jitkrittum, Harikrishna Narasimhan, Ankit Singh Rawat, Jeevesh Juneja et al.ICLR 2026 · 99 citations
- AutoJudge: Judge Decoding Without Manual AnnotationRoman Garipov, Fedor Velikonivtsev, Ivan Ermakov, Ruslan Svirschevski et al.NeurIPS 2025 · 14 citations
- LongSpec: Long-Context Lossless Speculative Decoding with Efficient Drafting and VerificationPenghui Yang, Cunxiao Du, Fengzhuo Zhang, Haonan Wang et al.ACL 2026 · 12 citations
- Gatekeeper: Improving Model Cascades Through Confidence TuningStephan Rabanser, Nathalie Rauschmayr, Achin Kulshrestha, Petra Poklukar et al.NeurIPS 2025 · 10 citations
Builds on20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 1,472 citations
- FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPUYing Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li et al.ICML 2023 · 683 citations
- EAGLE: Speculative Sampling Requires Rethinking Feature UncertaintyYuhui Li, Fangyun Wei, Chao Zhang, Hongyang ZhangICML 2024 · 424 citations
- Confident Adaptive Language ModelingTal Schuster, Adam Fisch, Jai Gupta, Mostafa Dehghani et al.NeurIPS 2022 · 394 citations
Related papers
- Cascade Speculative Drafting for Even Faster LLM InferenceZiyi Chen, Xiaocong Yang, Jiacheng Lin, Chenkai Sun et al.NeurIPS 2024 · 107 citations
- Language Model Cascades: Token-Level Uncertainty And BeyondNeha Gupta, Harikrishna Narasimhan, Wittawat Jitkrittum, Ankit Singh Rawat et al.ICLR 2024 · 119 citations
- CAS-Spec: Cascade Adaptive Self-Speculative Decoding for On-the-Fly Lossless Inference Acceleration of LLMsZhiyuan Ning, Jiawei Shao, Ruge Xu, Xinfei Guo et al.NeurIPS 2025 · 5 citations
- SpecPIM: Accelerating Speculative Inference on PIM-Enabled System via Architecture-Dataflow Co-ExplorationCong Li, Zhe Zhou, Size Zheng, Jiaxi Zhang et al.ASPLOS 2024 · 29 citations
- SpecExec: Massively Parallel Speculative Decoding For Interactive LLM Inference on Consumer DevicesRuslan Svirschevski, Avner May, Zhuoming Chen, Beidi Chen et al.NeurIPS 2024 · 70 citations
