AdaSpec: Adaptive Multilingual Speculative Decoding with Self-Synthesized Language-Aware Training and Vocabulary Simplification
Dinh-Truong Do, Nguyen-Khang Le, Le-Minh Nguyen
Abstract
Speculative decoding accelerates large language model (LLM) inference by using a lightweight drafter to propose multiple tokens, which are then verified in parallel by the base model. While effective in English, existing methods often struggle in multilingual scenarios due to static vocabularies and the lack of language-specific instruction data. To address these limitations, we present AdaSpec, a multilingual speculative decoding framework that dynamically adapts both the drafter and vocabulary at decoding time. AdaSpec generates language-specific instruction data using the LLM itself, enabling training of drafters for low-resource languages. It also constructs adaptive vocabularies tailored to each language's characteristics. In addition, we introduce Multi-SpecBench, a comprehensive multilingual benchmark covering seven languages and seven generation tasks, to evaluate multilingual speculative decoding performance. Extensive experiments show that AdaSpec achieves up to 2.3× speedup over the state-of-the-art method of EAGLE-2, even in English, demonstrating its effectiveness across diverse languages and tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on11
- Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding HeadsTianle Cai, Yuhong Li, Zhengyang Geng, Hongwu Peng et al.ICML 2024 · 669 citations
- EAGLE: Speculative Sampling Requires Rethinking Feature UncertaintyYuhui Li, Fangyun Wei, Chao Zhang, Hongyang ZhangICML 2024 · 424 citations
- Break the Sequential Dependency of LLM Inference Using Lookahead DecodingYichao Fu, Peter Bailis, Ion Stoica, Hao ZhangICML 2024 · 290 citations
- SpecInfer: Accelerating Large Language Model Serving with Tree-based Speculative Inference and VerificationXupeng Miao, Gabriele Oliaro, Zhihao Zhang, Xinhao Cheng et al.ASPLOS 2024 · 105 citations
- Turning Trash into Treasure: Accelerating Inference of Large Language Models with Token RecyclingXianzhen Luo, Yixuan Wang, Qingfu Zhu, Zhiming Zhang et al.ACL 2025 · 31 citations
Related papers
- UniSpec: Training-Free Speculative Decoding for Robust LLM Acceleration Across Languages and HardwareTruong Dinh Do, Nguyen-Khang Le, Le-Minh NguyenACL 2026
- Not-a-Bandit: Provably No-Regret Drafter Selection in Speculative Decoding for LLMsHongyi Liu, Jiaji Huang, Zhen Jia, Youngsuk Park et al.ICLR 2026 · 5 citations
- SAM Decoding: Speculative Decoding via Suffix AutomatonYuxuan Hu, Ke Wang, Xiaokang Zhang, Fanjin Zhang et al.ACL 2025
- PARD: Accelerating LLM Inference with Low‑Cost PARallel Draft Model AdaptationZihao An, Huajun Bai, Ziqiong Liu, Dong Li et al.ICLR 2026 · 28 citations
- NanoSpec: Accelerating Speculative Decoding using Minimalist In-Context VocabulariesZhiyang Chen, Daliang Xu, Yinyuan Zhang, Chenghua Wang et al.ICML 2026 · 1 citation
