Talon: Breaking the Synchronization Barrier in Speculative Decoding with Hybrid Model-based and Retrieve-based Drafting
Xiangxiang Gao, Weisheng Xie, Lixin, Xuwei Fang, Chen Hang, Changqun Li, Yuhan Lin, Xiaolong Xu
Abstract
Large Language Models face fundamental deployment challenges due to the computational demands of auto-regressive token-by-token generation. While speculative decoding has emerged as a promising acceleration technique through its draft-then-verify framework, current implementations suffer from two critical limitations: (1) mutual waiting problem caused by sequential dependencies between draft generation and verification phases, and (2) constrained token acceptance rates where retrieval-based drafting methods under-perform in general domains while models-based drafting approaches show reduced efficacy in knowledge-intensive scenarios. To address these challenges, we propose Talon, a novel parallel inference architecture featuring two key innovations: (1) a novel asynchronous execution paradigm that decouples draft generation from verification, effectively eliminating synchronization bottlenecks, and (2) an adaptive hybrid drafting strategy that dynamically combines model-based and retrieval-based approaches to improve token acceptance rates across diverse domains. Extensive evaluations across standard benchmarks (MT-Bench, HumanEval, GSM8K, Alpaca, CNN/DM) demonstrate Talon's exceptional performance, achieving 4.04x–6.52x acceleration across multiple model families including Vicuna, Deepseek, and LLaMA series. These results represent a significant advancement over existing speculative decoding methods (EAGLE 1-3, Hydra, Medusa, Lookahead, SPS, and PLD), establishing a new paradigm for speculative decoding.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 680ce13c-7095-4ba2-8abd-a4954bb8a6d9Builds on8
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 1,472 citations
- Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding HeadsTianle Cai, Yuhong Li, Zhengyang Geng, Hongwu Peng et al.ICML 2024 · 669 citations
- EAGLE: Speculative Sampling Requires Rethinking Feature UncertaintyYuhui Li, Fangyun Wei, Chao Zhang, Hongyang ZhangICML 2024 · 424 citations
- EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time TestYuhui Li, Fangyun Wei, Chao Zhang, Hongyang ZhangNeurIPS 2025 · 347 citations
- Break the Sequential Dependency of LLM Inference Using Lookahead DecodingYichao Fu, Peter Bailis, Ion Stoica, Hao ZhangICML 2024 · 290 citations
Related papers
- Falcon: Faster and Parallel Inference of Large Language Models Through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding TreeXiangxiang Gao, Weisheng Xie, Yiwei Xiang, Feng JiAAAI 2025 · 20 citations
- PARD: Accelerating LLM Inference with Low‑Cost PARallel Draft Model AdaptationZihao An, Huajun Bai, Ziqiong Liu, Dong Li et al.ICLR 2026 · 28 citations
- GRIFFIN: Effective Token Alignment for Faster Speculative DecodingShijing Hu, Jingyang Li, Xingyu Xie, Zhihui Lu et al.NeurIPS 2025 · 14 citations
- SpecBranch: Speculative Decoding via Hybrid Drafting and Rollback-Aware Branch ParallelismYuhao Shen, Junyi Shen, Quan Kong, Tianyu Liu et al.ICLR 2026 · 16 citations
- polybasic Speculative Decoding Through a Theoretical PerspectiveRuilin Wang, Huixia Li, Yuexiao Ma, Xiawu Zheng et al.ICML 2025
