SuffixDecoding: Extreme Speculative Decoding for Emerging AI Applications
Gabriele Oliaro, Zhihao Jia, Daniel F. Campos, Aurick Qiao
Abstract
Speculative decoding is widely adopted to reduce latency in large language model (LLM) inference by leveraging smaller draft models capable of handling diverse user tasks. However, emerging AI applications, such as LLM-based agents, present unique workload characteristics: instead of diverse independent requests, agentic frameworks typically submit repetitive inference requests, such as multi-agent pipelines performing similar subtasks or self-refinement loops iteratively enhancing outputs. These workloads result in long and highly predictable sequences, which current speculative decoding methods do not effectively exploit. To address this gap, we introduce SuffixDecoding, a novel method that utilizes efficient suffix trees to cache long token sequences from prompts and previous outputs. By adaptively speculating more tokens when acceptance likelihood is high and fewer when it is low, SuffixDecoding effectively exploits opportunities for longer speculations while conserving computation when those opportunities are limited. Evaluations on agentic benchmarks, including SWE-Bench and Text-to-SQL, demonstrate that SuffixDecoding achieves speedups of up to 5.3, outperforming state-of-the-art methods -- 2.8 faster than model-based approaches like EAGLE-2/3 and 1.9 faster than model-free approaches such as Token Recycling. SuffixDecoding is open-sourced at https://github.com/snowflakedb/ArcticInference
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 22620419-daa6-486e-a8cb-a105f41ea86aCited by top-tier papers9
- Seer: Online Context Learning for Fast Synchronous LLM Reinforcement LearningRuoyu Qin, Weiran He, Weixiao Huang, Yangkun Zhang et al.OSDI 2026 · 32 citations
- Speculative Speculative DecodingTanishq Kumar, Tri Dao, Avner MayICLR 2026 · 15 citations
- Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time ScalingShengyin Sun, Yiming Li, Xing Li, Yingzhao Lian et al.ICLR 2026 · 6 citations
- SelfJudge: Faster Speculative Decoding via Self-Supervised Judge VerificationKanghoon Yoon, Minsub Kim, Sungjae Lee, Joonhyung Lee et al.ICML 2026 · 4 citations
- Shift Parallelism: Low-Latency, High-Throughput LLM Inference for Dynamic WorkloadsMert Hidayetoglu, Aurick Qiao, Michael Wyatt, Jeff Rasley et al.ASPLOS 2026 · 3 citations
Builds on23
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan et al.NeurIPS 2023 · 4,972 citations
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao et al.ICLR 2024 · 2,082 citations
- SWE-agent: Agent-Computer Interfaces Enable Automated Software EngineeringJohn Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret et al.NeurIPS 2024 · 2,059 citations
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 1,472 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
Related papers
- LongSpec: Long-Context Lossless Speculative Decoding with Efficient Drafting and VerificationPenghui Yang, Cunxiao Du, Fengzhuo Zhang, Haonan Wang et al.ACL 2026 · 12 citations
- SAM Decoding: Speculative Decoding via Suffix AutomatonYuxuan Hu, Ke Wang, Xiaokang Zhang, Fanjin Zhang et al.ACL 2025
- Not-a-Bandit: Provably No-Regret Drafter Selection in Speculative Decoding for LLMsHongyi Liu, Jiaji Huang, Zhen Jia, Youngsuk Park et al.ICLR 2026 · 5 citations
- MagicDec: Breaking the Latency-Throughput Tradeoff for Long Context Generation with Speculative DecodingRanajoy Sadhukhan, Jian Chen, Zhuoming Chen, Vashisth Tiwari et al.ICLR 2025
- Inference-Cost-Aware Dynamic Tree Construction for Efficient Inference in Large Language ModelsYinrong Hong, Zhiquan Tan, Kai HuICLR 2026 · 2 citations
