SPECTRA: Faster Large Language Model Inference with Optimized Internal and External Speculation
Nguyen-Khang Le, Truong Dinh Do, Le-Minh Nguyen
Abstract
Inference with modern Large Language Models (LLMs) is both computationally expensive and time-consuming. Speculative decoding has emerged as a promising solution, but existing approaches face key limitations: training-based methods require a draft model that is challenging to obtain and lacks generalizability, while training-free methods offer limited speedup gains. In this work, we present SPECTRA, a novel framework for accelerating LLM inference without the need for additional training or modification to the original LLM. SPECTRA introduces two new techniques for efficiently utilizing internal and external speculation, each outperforming corresponding state-of-the-art (SOTA) methods independently. When combined, these techniques achieve up to a 4.08x speedup across various benchmarks and LLM architectures, significantly surpassing existing training-free approaches. The implementation of SPECTRA is publicly available.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3de3242e-261c-42d7-a3f5-8be913348a05Cited by top-tier papers2
- UniSpec: Training-Free Speculative Decoding for Robust LLM Acceleration Across Languages and HardwareTruong Dinh Do, Nguyen-Khang Le, Le-Minh NguyenACL 2026
- AdaSpec: Adaptive Multilingual Speculative Decoding with Self-Synthesized Language-Aware Training and Vocabulary SimplificationDinh-Truong Do, Nguyen-Khang Le, Le-Minh NguyenAAAI 2026
Builds on15
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 1,472 citations
- Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding HeadsTianle Cai, Yuhong Li, Zhengyang Geng, Hongwu Peng et al.ICML 2024 · 669 citations
- Sheared LLaMA: Accelerating Language Model Pre-training via Structured PruningMengzhou Xia, Tianyu Gao, Zhiyuan Zeng, Danqi ChenICLR 2024 · 453 citations
- Break the Sequential Dependency of LLM Inference Using Lookahead DecodingYichao Fu, Peter Bailis, Ion Stoica, Hao ZhangICML 2024 · 290 citations
Related papers
- SSSD: Simply-Scalable Speculative DecodingMichele Marzollo, Jiawei Zhuang, Niklas Roemer, Niklas Zwingenberger et al.ACL 2026 · 2 citations
- Turning Trash into Treasure: Accelerating Inference of Large Language Models with Token RecyclingXianzhen Luo, Yixuan Wang, Qingfu Zhu, Zhiming Zhang et al.ACL 2025 · 31 citations
- Draft& Verify: Lossless Large Language Model Acceleration via Self-Speculative DecodingJun Zhang, Jue Wang, Huan Li, Lidan Shou et al.ACL 2024
- Training-Free Loosely Speculative Decoding: Accepting Semantically Correct Drafts Beyond Exact MatchJinze Li, Yixing Xu, Guanchen Li, Shuo Yang et al.ICLR 2026 · 12 citations
- Ouroboros: Generating Longer Drafts Phrase by Phrase for Faster Speculative DecodingWeilin Zhao, Yuxiang Huang, Xu Han, Wang Xu et al.EMNLP 2024 · 4 citations
