Kangaroo: Lossless Self-Speculative Decoding for Accelerating LLMs via Double Early Exiting
Fangcheng Liu, Yehui Tang, Zhenhua Liu, Yunsheng Ni, Duyu Tang, Kai Han, Yunhe Wang
Abstract
Speculative decoding has demonstrated its effectiveness in accelerating the inference of large language models (LLMs) while maintaining an identical sampling distribution. However, the conventional approach of training separate draft model to achieve a satisfactory token acceptance rate can be costly and impractical. In this paper, we propose a novel self-speculative decoding framework Kangaroo with double early exiting strategy, which leverages the shallow sub-network and the LM Head of the well-trained target LLM to construct a self-drafting model. Then, the self-verification stage only requires computing the remaining layers over the early-exited hidden states in parallel. To bridge the representation gap between the sub-network and the full model, we train a lightweight and efficient adapter module on top of the sub-network. One significant challenge that comes with the proposed method is that the inference latency of the self-draft model may no longer be negligible compared to the big model. To boost the token acceptance rate while minimizing the latency of the self-drafting model, we introduce an additional early exiting mechanism for both single-sequence and the tree decoding scenarios. Specifically, we dynamically halt the small model’s subsequent prediction during the drafting phase once the confidence level for the current step falls below a certain threshold. This approach reduces unnecessary computations and improves overall efficiency. Extensive experiments on multiple benchmarks demonstrate our effectiveness, where Kangaroo achieves walltime speedups up to 2.04 × , outperforming Medusa-1 with 88.7% fewer additional parameters. The code for Kangaroo is available at https://github.com/Equationliu/Kangaroo .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers11
- FlashDLM: Accelerating Diffusion Language Model Inference via Efficient KV Caching and Guided DiffusionZhanqiu Hu, Jian Meng, Yash Akhauri, Mohamed S. Abdelfattah et al.ICLR 2026 · 56 citations
- SpecBranch: Speculative Decoding via Hybrid Drafting and Rollback-Aware Branch ParallelismYuhao Shen, Junyi Shen, Quan Kong, Tianyu Liu et al.ICLR 2026 · 16 citations
- STree: Speculative Tree Decoding for Hybrid State Space ModelsYangchao Wu, Zongyue Qin, Alex Wong, Stefano SoattoNeurIPS 2025 · 8 citations
- SpecForge: A Flexible and Efficient Open-Source Training Framework for Speculative DecodingShenggui Li, Chao Wang, YIKAI ZHU, Yubo Wang et al.ICML 2026 · 7 citations
- CAS-Spec: Cascade Adaptive Self-Speculative Decoding for On-the-Fly Lossless Inference Acceleration of LLMsZhiyuan Ning, Jiawei Shao, Ruge Xu, Xinfei Guo et al.NeurIPS 2025 · 5 citations
Builds on10
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language ModelsGuangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu et al.ICML 2023 · 1,493 citations
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 1,472 citations
- Confident Adaptive Language ModelingTal Schuster, Adam Fisch, Jai Gupta, Mostafa Dehghani et al.NeurIPS 2022 · 394 citations
- Break the Sequential Dependency of LLM Inference Using Lookahead DecodingYichao Fu, Peter Bailis, Ion Stoica, Hao ZhangICML 2024 · 290 citations
- DistillSpec: Improving Speculative Decoding via Knowledge DistillationYongchao Zhou, Kaifeng Lyu, Ankit Singh Rawat, Aditya Krishna Menon et al.ICLR 2024 · 143 citations
Related papers
- CLaSp: In-Context Layer Skip for Self-Speculative DecodingLongze Chen, Renke Shan, Huiming Wang, Lu Wang et al.ACL 2025
- Draft& Verify: Lossless Large Language Model Acceleration via Self-Speculative DecodingJun Zhang, Jue Wang, Huan Li, Lidan Shou et al.ACL 2024
- SWIFT: On-the-Fly Self-Speculative Decoding for LLM Inference AccelerationHeming Xia, Yongqi Li, Jun Zhang, Cunxiao Du et al.ICLR 2025
- Speculative Streaming: Efficient and Scalable Speculative Decoding with Multi-Stream AttentionNikhil Bhendawade, Irina Belousova, Qichen Fu, Henry Mason et al.EMNLP 2025
- Online Speculative DecodingXiaoxuan Liu, Lanxiang Hu, Peter Bailis, Alvin Cheung et al.ICML 2024 · 104 citations
