Kangaroo: Lossless Self-Speculative Decoding for Accelerating LLMs via Double Early Exiting
Fangcheng Liu, Yehui Tang, Zhenhua Liu, Yunsheng Ni, Duyu Tang, Kai Han, Yunhe Wang
摘要
Speculative decoding has demonstrated its effectiveness in accelerating the inference of large language models (LLMs) while maintaining an identical sampling distribution. However, the conventional approach of training separate draft model to achieve a satisfactory token acceptance rate can be costly and impractical. In this paper, we propose a novel self-speculative decoding framework Kangaroo with double early exiting strategy, which leverages the shallow sub-network and the LM Head of the well-trained target LLM to construct a self-drafting model. Then, the self-verification stage only requires computing the remaining layers over the early-exited hidden states in parallel. To bridge the representation gap between the sub-network and the full model, we train a lightweight and efficient adapter module on top of the sub-network. One significant challenge that comes with the proposed method is that the inference latency of the self-draft model may no longer be negligible compared to the big model. To boost the token acceptance rate while minimizing the latency of the self-drafting model, we introduce an additional early exiting mechanism for both single-sequence and the tree decoding scenarios. Specifically, we dynamically halt the small model’s subsequent prediction during the drafting phase once the confidence level for the current step falls below a certain threshold. This approach reduces unnecessary computations and improves overall efficiency. Extensive experiments on multiple benchmarks demonstrate our effectiveness, where Kangaroo achieves walltime speedups up to 2.04 × , outperforming Medusa-1 with 88.7% fewer additional parameters. The code for Kangaroo is available at https://github.com/Equationliu/Kangaroo .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- FlashDLM: Accelerating Diffusion Language Model Inference via Efficient KV Caching and Guided DiffusionZhanqiu Hu, Jian Meng, Yash Akhauri, Mohamed S. Abdelfattah 等ICLR 2026 · 被引用 56 次
- SpecBranch: Speculative Decoding via Hybrid Drafting and Rollback-Aware Branch ParallelismYuhao Shen, Junyi Shen, Quan Kong, Tianyu Liu 等ICLR 2026 · 被引用 16 次
- STree: Speculative Tree Decoding for Hybrid State Space ModelsYangchao Wu, Zongyue Qin, Alex Wong, Stefano SoattoNeurIPS 2025 · 被引用 8 次
- SpecForge: A Flexible and Efficient Open-Source Training Framework for Speculative DecodingShenggui Li, Chao Wang, YIKAI ZHU, Yubo Wang 等ICML 2026 · 被引用 7 次
- CAS-Spec: Cascade Adaptive Self-Speculative Decoding for On-the-Fly Lossless Inference Acceleration of LLMsZhiyuan Ning, Jiawei Shao, Ruge Xu, Xinfei Guo 等NeurIPS 2025 · 被引用 5 次
它引用的顶会 Paper10
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language ModelsGuangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu 等ICML 2023 · 被引用 1,493 次
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 被引用 1,472 次
- Confident Adaptive Language ModelingTal Schuster, Adam Fisch, Jai Gupta, Mostafa Dehghani 等NeurIPS 2022 · 被引用 394 次
- Break the Sequential Dependency of LLM Inference Using Lookahead DecodingYichao Fu, Peter Bailis, Ion Stoica, Hao ZhangICML 2024 · 被引用 290 次
- DistillSpec: Improving Speculative Decoding via Knowledge DistillationYongchao Zhou, Kaifeng Lyu, Ankit Singh Rawat, Aditya Krishna Menon 等ICLR 2024 · 被引用 143 次
相关 Paper
- CLaSp: In-Context Layer Skip for Self-Speculative DecodingLongze Chen, Renke Shan, Huiming Wang, Lu Wang 等ACL 2025
- Draft& Verify: Lossless Large Language Model Acceleration via Self-Speculative DecodingJun Zhang, Jue Wang, Huan Li, Lidan Shou 等ACL 2024
- SWIFT: On-the-Fly Self-Speculative Decoding for LLM Inference AccelerationHeming Xia, Yongqi Li, Jun Zhang, Cunxiao Du 等ICLR 2025
- Speculative Streaming: Efficient and Scalable Speculative Decoding with Multi-Stream AttentionNikhil Bhendawade, Irina Belousova, Qichen Fu, Henry Mason 等EMNLP 2025
- Online Speculative DecodingXiaoxuan Liu, Lanxiang Hu, Peter Bailis, Alvin Cheung 等ICML 2024 · 被引用 104 次
