Distributed Speculative Inference (DSI): Speculation Parallelism for Provably Faster Lossless Language Model Inference
Nadav Timor, Jonathan Mamou, Daniel Korat, Moshe Berchansky, Oren Pereg, Moshe Wasserblat, Tomer Galanti, Michal Gordon-Kiwkowitz, David Harel
2025年份
6顶会引用
摘要
This paper introduces distributed speculative inference (DSI), a novel inference algorithm that is provably faster than speculative inference (SI) (
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- SpecEdge: Scalable Edge-Assisted Serving Framework for Interactive LLMsJinwoo Park, Seunggeun Cho, Dongsu HanNeurIPS 2025 · 被引用 20 次
- SpecBranch: Speculative Decoding via Hybrid Drafting and Rollback-Aware Branch ParallelismYuhao Shen, Junyi Shen, Quan Kong, Tianyu Liu 等ICLR 2026 · 被引用 16 次
- Speculative Speculative DecodingTanishq Kumar, Tri Dao, Avner MayICLR 2026 · 被引用 15 次
- Double: Breaking the Acceleration Limit via Double Retrieval Speculative ParallelismYuhao Shen, Tianyu Liu, Junyi Shen, Jinyang Wu 等ACL 2026 · 被引用 12 次
- MineDraft: A Framework for Batch Parallel Speculative DecodingZhenwei Tang, Arun Verma, Zijian Zhou, Zhaoxuan Wu 等ICML 2026
它引用的顶会 Paper29
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han 等ICLR 2024 · 被引用 1,714 次
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 被引用 1,472 次
相关 Paper
- Accelerated Speculative Sampling Based on Tree Monte CarloZhengmian Hu, Heng HuangICML 2024 · 被引用 16 次
- MPCFORMER: Fast, Performant and Provate Transformer Inference with MPCDacheng Li, Hongyi Wang, Rulin Shao, Han Guo 等ICLR 2023
- SpecStega: Provably Secure Linguistic Steganography Based on Speculative Sampling in Asymmetric Resource ScenariosJun Jiang, Kejiang Chen, Yuang Qi, Jiawei Zhao 等CCS 2026
- GoodSpeed: Optimizing Fair Goodput with Adaptive Speculative Decoding in Distributed Edge InferencePhuong Tran, Tzu-Hao Liu, Long Tan Le, Tung-Anh Nguyen 等INFOCOM 2026
- SpecInfer: Accelerating Large Language Model Serving with Tree-based Speculative Inference and VerificationXupeng Miao, Gabriele Oliaro, Zhihao Zhang, Xinhao Cheng 等ASPLOS 2024 · 被引用 105 次
