Mixture of Attentions For Speculative Decoding
Matthieu Zimmer, Milan Gritta, Gerasimos Lampouras, Haitham Bou-Ammar, Jun Wang
摘要
The growth in the number of parameters of Large Language Models (LLMs) has led to a significant surge in computational requirements, making them challenging and costly to deploy. Speculative decoding (SD) leverages smaller models to efficiently propose future tokens, which are then verified by the LLM in parallel. Small models that utilise activations from the LLM currently achieve the fastest decoding speeds. However, we identify several limitations of SD models including the lack of on-policyness during training and partial observability. To address these shortcomings, we propose a more grounded architecture for small models by introducing a Mixture of Attentions for SD. Our novel architecture can be applied in two scenarios: a conventional single device deployment and a novel client-server deployment where the small model is hosted on a consumer device and the LLM on a server. In a single-device scenario, we demonstrate state-ofthe-art speedups improving EAGLE-2 by 9.5% and its acceptance length by 25%. In a client-server setting, our experiments demonstrate: 1) state-of-the-art latencies with minimal calls to the server for different network conditions, and 2) in the event of a complete disconnection, our approach can maintain higher accuracy compared to other SD methods and demonstrates advantages over API calls to LLMs, which would otherwise be unable to continue the generation process.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Scalable Power Sampling: Unlocking Efficient, Training-Free Reasoning for LLMs via Distribution SharpeningXiaotong Ji, Rasul Tutunov, Matthieu Zimmer, Haitham Bou AmmarICML 2026 · 被引用 17 次
- ConfSpec: Efficient Step-Level Speculative Reasoning via Confidence-Gated VerificationSiran Liu, Zane Cao, Yongchao HeACL 2026 · 被引用 3 次
- HeteroSpec: Leveraging Contextual Heterogeneity for Efficient Speculative DecodingSiran Liu, Yang Ye, Qianchao Zhu, Zane Cao 等ACL 2026 · 被引用 2 次
- Gumiho: A Hybrid Architecture to Prioritize Early Tokens in Speculative DecodingJinze Li, Yixing Xu, Haiduo Huang, Xuanwu Yin 等ICML 2025
- Steering Pretrained Drafters During Speculative DecodingFrédéric Berdoz, Peer Rheinboldt, Roger WattenhoferAAAI 2026
它引用的顶会 Paper16
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 被引用 1,472 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- Break the Sequential Dependency of LLM Inference Using Lookahead DecodingYichao Fu, Peter Bailis, Ion Stoica, Hao ZhangICML 2024 · 被引用 290 次
- Better & Faster Large Language Models via Multi-token PredictionFabian Gloeckle, Badr Youbi Idrissi, Baptiste Rozière, David Lopez-Paz 等ICML 2024 · 被引用 286 次
相关 Paper
- ConFu: Contemplate the Future for Better Speculative SamplingZongyue Qin, Raghavv Goel, Mukul Gagrani, Risheek Garrepalli 等ICML 2026
- Speculative Streaming: Efficient and Scalable Speculative Decoding with Multi-Stream AttentionNikhil Bhendawade, Irina Belousova, Qichen Fu, Henry Mason 等EMNLP 2025
- RepSpec: Structural Re-parameterized Draft Model Training for Speculative DecodingFeiye Huo, Jianchao Tan, Jiahao Liu, Zixu Jiang 等ICLR 2026
- EAGLE: Speculative Sampling Requires Rethinking Feature UncertaintyYuhui Li, Fangyun Wei, Chao Zhang, Hongyang ZhangICML 2024 · 被引用 424 次
- Not-a-Bandit: Provably No-Regret Drafter Selection in Speculative Decoding for LLMsHongyi Liu, Jiaji Huang, Zhen Jia, Youngsuk Park 等ICLR 2026 · 被引用 5 次
