PARD: Accelerating LLM Inference with Low‑Cost PARallel Draft Model Adaptation
Zihao An, Huajun Bai, Ziqiong Liu, Dong Li, Emad Barsoum
摘要
The autoregressive nature of large language models (LLMs) fundamentally limits inference speed, as each forward pass generates only a single token and is often bottlenecked by memory bandwidth. Speculative decoding has emerged as a promising solution, adopting a draft-then-verify strategy to accelerate token generation. While the EAGLE series achieves strong acceleration, its requirement of training a separate draft head for each target model introduces substantial adaptation costs. In this work, we propose PARD (PARallel Draft), a novel speculative decoding method featuring target-independence and parallel token prediction. Specifically, PARD enables a single draft model to be applied across an entire family of target models without requiring separate training for each variant, thereby minimizing adaptation costs. Meanwhile, PARD substantially accelerates inference by predicting multiple future tokens within a single forward pass of the draft phase. To further reduce the training adaptation cost of PARD, we propose a COnditional Drop-token (COD) mechanism based on the integrity of prefix key-value states, enabling autoregressive draft models to be adapted into parallel draft models at low-cost. Our experiments show that the proposed COD method improves draft model training efficiency by 3× compared with traditional masked prediction training. On the vLLM inference framework, PARD achieves up to 3.67× speedup on LLaMA3.1-8B, reaching 264.88 tokens per second, which is 1.15× faster than EAGLE-3. Our code is available at https://github.com/AMD-AIG-AIMA/PARD .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- SPEED-Bench: A Unified and Diverse Benchmark for Speculative DecodingTalor Abramovich, Maor Ashkenazi, Izzy Putterman, Benjamin Chislett 等ICML 2026
- InstEmb: Instruction-Following Embeddings through Glimpses of the FutureTianhao Gao, Jun Fang, Xiaohui Zhang, Zhiyuan Liu 等ICML 2026
它引用的顶会 Paper15
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han 等ICLR 2024 · 被引用 1,714 次
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language ModelsGuangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu 等ICML 2023 · 被引用 1,493 次
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 被引用 1,472 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
相关 Paper
- ConFu: Contemplate the Future for Better Speculative SamplingZongyue Qin, Raghavv Goel, Mukul Gagrani, Risheek Garrepalli 等ICML 2026
- HCSpec: Two-Tier Horizontal Cascade Speculative Decoding for High-Efficiency Large Language Model InferenceYizhou Zhang, Siming Chen, Hao Ye, Erhu FengACL 2026
- Training-Free Loosely Speculative Decoding: Accepting Semantically Correct Drafts Beyond Exact MatchJinze Li, Yixing Xu, Guanchen Li, Shuo Yang 等ICLR 2026 · 被引用 12 次
- CORAL: Learning Consistent Representations across Multi-step Training with Lighter Speculative DrafterYepeng Weng, Dianwen Mei, Huishi Qiu, Xujie Chen 等ACL 2025 · 被引用 6 次
- Talon: Breaking the Synchronization Barrier in Speculative Decoding with Hybrid Model-based and Retrieve-based DraftingXiangxiang Gao, Weisheng Xie, Lixin, Xuwei Fang 等AAAI 2026
