KnapSpec: Self-Speculative Decoding via Adaptive Layer Selection as a Knapsack Problem
Seongjin Cha, Gyuwan Kim, Dongsu Han, Tao Yang, Insu Han
摘要
Self-speculative decoding (SSD) accelerates LLM inference by skipping layers to create an efficient draft model, yet existing methods often rely on static heuristics that ignore the dynamic computational overhead of attention in long-context scenarios. We propose KnapSpec, a training-free framework that reformulates draft model selection as a knapsack problem to maximize tokensper-time throughput. By decoupling Attention and MLP layers and modeling their hardwarespecific latencies as functions of context length, KnapSpec adaptively identifies optimal draft configurations on the fly via a parallel dynamic programming algorithm. Furthermore, we provide the first rigorous theoretical analysis establishing cosine similarity between hidden states as a mathematically sound proxy for the token acceptance rate. This foundation allows our method to maintain high drafting faithfulness while navigating the shifting bottlenecks of real-world hardware. Our experiments on Qwen3 and Llama3 demonstrate that KnapSpec consistently outperforms state-of-the-art SSD baselines, achieving up to 1.47× wall-clock speedup across various benchmarks. Our plug-and-play approach ensures high-speed inference for long sequences without requiring additional training or compromising the target model's output distribution. Our official implementation is available at https://github. com/kaist-flexml-lab/knapspec.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question AnsweringPan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu 等NeurIPS 2022 · 被引用 2,727 次
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris 等UIST 2023 · 被引用 1,882 次
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 被引用 1,472 次
- Compressive Transformers for Long-Range Sequence ModellingJack W. Rae, Anna Potapenko, Siddhant M. Jayakumar, Chloe Hillier 等ICLR 2020 · 被引用 833 次
相关 Paper
- SWIFT: On-the-Fly Self-Speculative Decoding for LLM Inference AccelerationHeming Xia, Yongqi Li, Jun Zhang, Cunxiao Du 等ICLR 2025
- CLaSp: In-Context Layer Skip for Self-Speculative DecodingLongze Chen, Renke Shan, Huiming Wang, Lu Wang 等ACL 2025
- Vegas: Self-Speculative Decoding with Verification-Guided Sparse AttentionYikang Yue, Yuqi Xue, Jian HuangICML 2026 · 被引用 2 次
- LongSpec: Long-Context Lossless Speculative Decoding with Efficient Drafting and VerificationPenghui Yang, Cunxiao Du, Fengzhuo Zhang, Haonan Wang 等ACL 2026 · 被引用 12 次
- Training-Free Loosely Speculative Decoding: Accepting Semantically Correct Drafts Beyond Exact MatchJinze Li, Yixing Xu, Guanchen Li, Shuo Yang 等ICLR 2026 · 被引用 12 次
