Distributed Speculative Inference (DSI): Speculation Parallelism for Provably Faster Lossless Language Model Inference
Nadav Timor, Jonathan Mamou, Daniel Korat, Moshe Berchansky, Oren Pereg, Moshe Wasserblat, Tomer Galanti, Michal Gordon-Kiwkowitz, David Harel
2025Year
6Top-tier citations
Abstract
This paper introduces distributed speculative inference (DSI), a novel inference algorithm that is provably faster than speculative inference (SI) (
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- SpecEdge: Scalable Edge-Assisted Serving Framework for Interactive LLMsJinwoo Park, Seunggeun Cho, Dongsu HanNeurIPS 2025 · 20 citations
- SpecBranch: Speculative Decoding via Hybrid Drafting and Rollback-Aware Branch ParallelismYuhao Shen, Junyi Shen, Quan Kong, Tianyu Liu et al.ICLR 2026 · 16 citations
- Speculative Speculative DecodingTanishq Kumar, Tri Dao, Avner MayICLR 2026 · 15 citations
- Double: Breaking the Acceleration Limit via Double Retrieval Speculative ParallelismYuhao Shen, Tianyu Liu, Junyi Shen, Jinyang Wu et al.ACL 2026 · 12 citations
- MineDraft: A Framework for Batch Parallel Speculative DecodingZhenwei Tang, Arun Verma, Zijian Zhou, Zhaoxuan Wu et al.ICML 2026
Builds on29
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 1,472 citations
Related papers
- Accelerated Speculative Sampling Based on Tree Monte CarloZhengmian Hu, Heng HuangICML 2024 · 16 citations
- MPCFORMER: Fast, Performant and Provate Transformer Inference with MPCDacheng Li, Hongyi Wang, Rulin Shao, Han Guo et al.ICLR 2023
- SpecStega: Provably Secure Linguistic Steganography Based on Speculative Sampling in Asymmetric Resource ScenariosJun Jiang, Kejiang Chen, Yuang Qi, Jiawei Zhao et al.CCS 2026
- GoodSpeed: Optimizing Fair Goodput with Adaptive Speculative Decoding in Distributed Edge InferencePhuong Tran, Tzu-Hao Liu, Long Tan Le, Tung-Anh Nguyen et al.INFOCOM 2026
- SpecInfer: Accelerating Large Language Model Serving with Tree-based Speculative Inference and VerificationXupeng Miao, Gabriele Oliaro, Zhihao Zhang, Xinhao Cheng et al.ASPLOS 2024 · 105 citations
