Lune

ACL2026Top-tier venue

EDSD: Entropy-Driven Design for Faster Speculative Decoding

Longkai Cheng, Ximing Wang, Jiangcai Zhu, Kailai Shao, Chao Chen, Haixiang Hu

2026Year

Abstract

Speculative decoding has emerged as a promising paradigm for accelerating large language model inference by leveraging a lightweight draft model to generate multiple candidate tokens. However, existing methods often incur substantial training overhead to mitigate information misalignment between autoregressive draft model training and decoding. To address this challenge, we propose EDSD, an Entropy-Driven Speculative Decoding framework that uses entropy as a unified, interpretable signal for both draft model training and architectural design. EDSD drives the draft model to progressively align with the target model in an easy-to-hard manner while establishing tokenlevel alignment as a dominant design principle. Extensive experiments on seven LLMs demonstrate that EDSD improves training efficiency by 24.8%, increases the average acceptance length by 4.0%, and achieves a 4.1% speedup compared to state-of-the-art methods. Furthermore, EDSD improves robustness to system prompt variations by more than 5×. Our findings establish entropy-driven alignment as an effective and principled foundation for efficient speculative decoding. We make our draft model weights available at https://github.com/KerwinKai/EDSD .

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext dd17f4dc-a929-4efe-b260-24d31be4af47

Builds on25

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines