LongSpec: Long-Context Lossless Speculative Decoding with Efficient Drafting and Verification
Penghui Yang, Cunxiao Du, Fengzhuo Zhang, Haonan Wang, Tianyu Pang, Chao Du, Bo An
摘要
As Large Language Models (LLMs) can now process extremely long contexts, efficient inference over these extended inputs has become increasingly important, especially for emerging applications like LLM agents that highly depend on this capability. Speculative decoding (SD) offers a promising lossless acceleration technique compared to lossy alternatives such as quantization and model cascades. However, most state-of-the-art SD methods are trained on short texts (typically fewer than 4k tokens), making them unsuitable for long-context scenarios. Specifically, adapting these methods to long contexts presents three key challenges: (1) the excessive memory demands posed by draft models due to large Key-Value (KV) cache; (2) performance degradation resulting from the mismatch between short-context training and long-context inference; and (3) inefficiencies in tree attention mechanisms when managing long token sequences. This work introduces LongSpec, a framework that addresses these challenges through three core innovations: a memory-efficient draft model with a constant-sized KV cache; novel position indices that mitigate the training-inference mismatch; and an attention aggregation strategy that combines fast prefix computation with standard tree attention to enable efficient decoding. Experimental results confirm the effectiveness of LongSpec, achieving up to a 3.26x speedup over strong Flash Attention baselines across five long-context understanding datasets, as well as a 2.25x reduction in wall-clock time on the AIME24 long reasoning task with the QwQ model, demonstrating significant latency improvements for long-context applications. The code is available at https://github.com/sail-sg/LongSpec.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM InferenceYuan Feng, Junlin Lv, Yukun Cao, Xike Xie 等NeurIPS 2025 · 被引用 256 次
- Speculative Speculative DecodingTanishq Kumar, Tri Dao, Avner MayICLR 2026 · 被引用 15 次
- OmniDraft: A cross-vocabulary, online adaptive drafter for on-device speculative decodingRamchalam Kinattinkara Ramakrishnan, Zhaocong Yuan, Jay Zhuo, Chen Feng 等NeurIPS 2025 · 被引用 6 次
- ForesightKV: Optimizing KV Cache Eviction for Reasoning Models by Learning Long-Term ContributionZican Dong, Peiyu Liu, Junyi Li, Zhipeng Chen 等ICML 2026
- Dustin: Draft-Augmented Sparse Verification for Efficient Long-Context Generation with Speculative DecodingWen-Hung Lee, Jian-Jia Chen, Xiaolin Lin, Pei-Shuo Wang 等ICML 2026
它引用的顶会 Paper29
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han 等ICLR 2024 · 被引用 1,714 次
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 被引用 1,472 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding HeadsTianle Cai, Yuhong Li, Zhengyang Geng, Hongwu Peng 等ICML 2024 · 被引用 669 次
- YaRN: Efficient Context Window Extension of Large Language ModelsBowen Peng, Jeffrey Quesnelle, Honglu Fan, Enrico ShippoleICLR 2024 · 被引用 508 次
相关 Paper
- QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV CacheRishabh Tiwari, Haocheng Xi, Aditya Tomar, Coleman Richard Charles Hooper 等ICML 2025
- RAPID: Long-Context Inference with Retrieval-Augmented Speculative DecodingGuanzheng Chen, Qilong Feng, Jinjie Ni, Xin Li 等ICML 2025
- MagicDec: Breaking the Latency-Throughput Tradeoff for Long Context Generation with Speculative DecodingRanajoy Sadhukhan, Jian Chen, Zhuoming Chen, Vashisth Tiwari 等ICLR 2025
- SuffixDecoding: Extreme Speculative Decoding for Emerging AI ApplicationsGabriele Oliaro, Zhihao Jia, Daniel F. Campos, Aurick QiaoNeurIPS 2025 · 被引用 34 次
- KnapSpec: Self-Speculative Decoding via Adaptive Layer Selection as a Knapsack ProblemSeongjin Cha, Gyuwan Kim, Dongsu Han, Tao Yang 等ICML 2026
