PARD: Accelerating LLM Inference with Low‑Cost PARallel Draft Model Adaptation
Zihao An, Huajun Bai, Ziqiong Liu, Dong Li, Emad Barsoum
Abstract
The autoregressive nature of large language models (LLMs) fundamentally limits inference speed, as each forward pass generates only a single token and is often bottlenecked by memory bandwidth. Speculative decoding has emerged as a promising solution, adopting a draft-then-verify strategy to accelerate token generation. While the EAGLE series achieves strong acceleration, its requirement of training a separate draft head for each target model introduces substantial adaptation costs. In this work, we propose PARD (PARallel Draft), a novel speculative decoding method featuring target-independence and parallel token prediction. Specifically, PARD enables a single draft model to be applied across an entire family of target models without requiring separate training for each variant, thereby minimizing adaptation costs. Meanwhile, PARD substantially accelerates inference by predicting multiple future tokens within a single forward pass of the draft phase. To further reduce the training adaptation cost of PARD, we propose a COnditional Drop-token (COD) mechanism based on the integrity of prefix key-value states, enabling autoregressive draft models to be adapted into parallel draft models at low-cost. Our experiments show that the proposed COD method improves draft model training efficiency by 3× compared with traditional masked prediction training. On the vLLM inference framework, PARD achieves up to 3.67× speedup on LLaMA3.1-8B, reaching 264.88 tokens per second, which is 1.15× faster than EAGLE-3. Our code is available at https://github.com/AMD-AIG-AIMA/PARD .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 179f9647-e264-4241-bcdb-a62939633e0fCited by top-tier papers2
- SPEED-Bench: A Unified and Diverse Benchmark for Speculative DecodingTalor Abramovich, Maor Ashkenazi, Izzy Putterman, Benjamin Chislett et al.ICML 2026
- InstEmb: Instruction-Following Embeddings through Glimpses of the FutureTianhao Gao, Jun Fang, Xiaohui Zhang, Zhiyuan Liu et al.ICML 2026
Builds on15
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language ModelsGuangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu et al.ICML 2023 · 1,493 citations
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 1,472 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
Related papers
- ConFu: Contemplate the Future for Better Speculative SamplingZongyue Qin, Raghavv Goel, Mukul Gagrani, Risheek Garrepalli et al.ICML 2026
- HCSpec: Two-Tier Horizontal Cascade Speculative Decoding for High-Efficiency Large Language Model InferenceYizhou Zhang, Siming Chen, Hao Ye, Erhu FengACL 2026
- Training-Free Loosely Speculative Decoding: Accepting Semantically Correct Drafts Beyond Exact MatchJinze Li, Yixing Xu, Guanchen Li, Shuo Yang et al.ICLR 2026 · 12 citations
- CORAL: Learning Consistent Representations across Multi-step Training with Lighter Speculative DrafterYepeng Weng, Dianwen Mei, Huishi Qiu, Xujie Chen et al.ACL 2025 · 6 citations
- Talon: Breaking the Synchronization Barrier in Speculative Decoding with Hybrid Model-based and Retrieve-based DraftingXiangxiang Gao, Weisheng Xie, Lixin, Xuwei Fang et al.AAAI 2026
