TETRIS: Optimal Draft Token Selection for Batch Speculative Decoding
Zhaoxuan Wu, Zijian Zhou, Arun Verma, Alok Prakash, Daniela Rus, Bryan Kian Hsiang Low
Abstract
We propose TETRIS, a novel method that optimizes the total throughput of batch speculative decoding in multi-request settings. Unlike existing methods that optimize for a single request or a group of requests as a whole, TETRIS actively selects the most promising draft tokens (for every request in a batch) to be accepted when verified in parallel, resulting in fewer rejected tokens and hence less wasted computing resources. Such an effective resource utilization to achieve fast inference in large language models (LLMs) is especially important to service providers with limited inference capacity. Compared to baseline speculative decoding, TETRIS yields a consistently higher acceptance rate and more effective utilization of the limited inference capacity. We show theoretically and empirically that TETRIS outperforms baseline speculative decoding and existing methods that dynamically select draft tokens, leading to a more efficient batch inference in LLMs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6f9850fe-e962-4469-8672-81bdc05a328eCited by top-tier papers2
- ECHO: Elastic Speculative Decoding with Sparse Gating for High-Concurrency ScenariosXinyi Hu, Yuhao Shen, Zhang Baolin, Hengxin Zhang et al.ICML 2026 · 7 citations
- Inference-Cost-Aware Dynamic Tree Construction for Efficient Inference in Large Language ModelsYinrong Hong, Zhiquan Tan, Kai HuICLR 2026 · 2 citations
Builds on16
- SGLang: Efficient Execution of Structured Language Model ProgramsLianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun et al.NeurIPS 2024 · 1,586 citations
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 1,472 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding HeadsTianle Cai, Yuhong Li, Zhengyang Geng, Hongwu Peng et al.ICML 2024 · 669 citations
- EAGLE: Speculative Sampling Requires Rethinking Feature UncertaintyYuhui Li, Fangyun Wei, Chao Zhang, Hongyang ZhangICML 2024 · 424 citations
Related papers
- Tetris: Efficient Long-context LLM Serving with Chunkwise Dynamic Sequence ParallelismCong Li, Yuzhe Yang, Xuegui Zheng, Qifan Yang et al.ISCA 2026
- Block Verification Accelerates Speculative DecodingZiteng Sun, Uri Mendlovic, Yaniv Leviathan, Asaf Aharoni et al.ICLR 2025
- SPIN: Accelerating Large Language Model Inference with Heterogeneous Speculative ModelsFahao Chen, Peng Li, Tom H. Luan, Zhou Su et al.INFOCOM 2025 · 10 citations
- Towards Efficient LLM Inference via Collective and Adaptive Speculative DecodingSiqi Wang, Hailong Yang, Xuezhu Wang, Tongxuan Liu et al.SC 2025 · 3 citations
- Speculative Decoding with CTC-based Draft Model for LLM Inference AccelerationZhuofan Wen, Shangtong Gui, Yang FengNeurIPS 2024 · 19 citations
