TokenTiming: A Dynamic Alignment Method for Universal Speculative Decoding Model Pairs
Sibo Xiao, Jinyuan Fu, Zhongle Xie, Lidan Shou
Abstract
Accelerating the inference of large language models (LLMs) has been a critical challenge in generative AI. Speculative decoding (SD) substantially improves LLM inference efficiency. However, its utility is limited by a fundamental constraint: the draft and target models must share the same vocabulary, thus limiting the herd of available draft models and often necessitating the training of a new model from scratch. Inspired by Dynamic Time Warping (DTW), a classic algorithm for aligning time series, we propose the algorithm TokenTiming for universal speculative decoding. It operates by re-encoding the draft token sequence to get a new target token sequence, and then uses DTW to build a mapping to transfer the probability distributions for speculative sampling. Benefiting from this, our method accommodates mismatched vocabularies and works with any off-the-shelf models without re-training and modification. We conduct comprehensive experiments on various tasks, demonstrating 1.57× speedup. This work enables a universal approach for draft model selection, making SD a more versatile and practical tool for LLM acceleration. The code is available at the link.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a9ab6664-cced-4f69-8f99-b16c5f9d2c4cBuilds on9
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 1,472 citations
- Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding HeadsTianle Cai, Yuhong Li, Zhengyang Geng, Hongwu Peng et al.ICML 2024 · 669 citations
- EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time TestYuhui Li, Fangyun Wei, Chao Zhang, Hongyang ZhangNeurIPS 2025 · 347 citations
- Break the Sequential Dependency of LLM Inference Using Lookahead DecodingYichao Fu, Peter Bailis, Ion Stoica, Hao ZhangICML 2024 · 290 citations
- Neural Machine Translation with Byte-Level SubwordsChanghan Wang, Kyunghyun Cho, Jiatao GuAAAI 2020 · 213 citations
Related papers
- Accelerating LLM Inference with Lossless Speculative Decoding Algorithms for Heterogeneous VocabulariesNadav Timor, Jonathan Mamou, Daniel Korat, Moshe Berchansky et al.ICML 2025
- A Theoretical Perspective for Speculative Decoding AlgorithmMing Yin, Minshuo Chen, Kaixuan Huang, Mengdi WangNeurIPS 2024 · 36 citations
- Speculative Decoding with CTC-based Draft Model for LLM Inference AccelerationZhuofan Wen, Shangtong Gui, Yang FengNeurIPS 2024 · 19 citations
- SpecPIM: Accelerating Speculative Inference on PIM-Enabled System via Architecture-Dataflow Co-ExplorationCong Li, Zhe Zhou, Size Zheng, Jiaxi Zhang et al.ASPLOS 2024 · 29 citations
- DistillSpec: Improving Speculative Decoding via Knowledge DistillationYongchao Zhou, Kaifeng Lyu, Ankit Singh Rawat, Aditya Krishna Menon et al.ICLR 2024 · 143 citations
