SC2024Top-tier venue
PipeInfer: Accelerating LLM Inference using Asynchronous Pipelined Speculation
Branden Butler, Sixing Yu, Arya Mazaheri, Ali Jannesari
Abstract
Inference of Large Language Models (LLMs) across computer clusters has become a focal point of research in recent times, with many acceleration techniques taking inspiration from CPU speculative execution. These techniques reduce bottlenecks associated with memory bandwidth, but also increase end-to-end latency per inference run, requiring high speculation acceptance rates to improve performance. Combined with a variable rate of acceptance across tasks, speculative inference techniques can result in reduced performance. Additionally, pipeline-parallel designs require many user requests to maintain maximum utilization. As a remedy, we propose PipeInfer, a pipelined speculative acceleration technique to reduce inter-token latency and improve system utilization for single-request scenarios while also improving tolerance to low speculation acceptance rates and low-bandwidth interconnects. PipeInfer exhibits up to a improvement in generation speed over standard speculative inference. PipeInfer achieves its improvement through Continuous Asynchronous Speculation and Early Inference Cancellation, the former improving latency and generation speed by running single-token inference simultaneously with several speculative runs, while the latter improves speed and latency by skipping the computation of invalidated runs, even in the middle of inference.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 000564ce-e7f8-43c1-a250-a1159496e72aCited by top-tier papers4
- HydraServe: Minimizing Cold Start Latency for Serverless LLM Serving in Public CloudsChiheng Lou, Sheng Qi, Chao Jin, Dapeng Nie et al.NSDI 2026 · 22 citations
- MaverIQ: Fingerprint-Guided Extrapolation and Fragmentation-Aware Layering for Intent-Based LLM ServingDimitrios Liakopoulos, Prasoon Sinha, Tianrui Hu, Myungjin Lee et al.SC 2025 · 2 citations
- TokenSwift: Lossless Acceleration of Ultra Long Sequence GenerationTong Wu, Junzhe Shen, Zixia Jia, Yuxuan Wang et al.ICML 2025
- EditFlow: Benchmarking and Optimizing Code Edit Recommendation Systems via Reconstruction of Developer FlowsChenyan Liu, Yun Lin, Jiaxin Chang, Jiawei Liu et al.OOPSLA 2026
Builds on7
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- SparseGPT: Massive Language Models Can be Accurately Pruned in One-ShotElias Frantar, Dan AlistarhICML 2023 · 1,240 citations
- QuIP: 2-Bit Quantization of Large Language Models With GuaranteesJerry Chee, Yaohui Cai, Volodymyr Kuleshov, Christopher De SaNeurIPS 2023 · 503 citations
- Confident Adaptive Language ModelingTal Schuster, Adam Fisch, Jai Gupta, Mostafa Dehghani et al.NeurIPS 2022 · 394 citations
- SqueezeLLM: Dense-and-Sparse QuantizationSehoon Kim, Coleman Hooper, Amir Gholami, Zhen Dong et al.ICML 2024 · 306 citations
Related papers
- SpecInfer: Accelerating Large Language Model Serving with Tree-based Speculative Inference and VerificationXupeng Miao, Gabriele Oliaro, Zhihao Zhang, Xinhao Cheng et al.ASPLOS 2024 · 105 citations
- Speculative Decoding with CTC-based Draft Model for LLM Inference AccelerationZhuofan Wen, Shangtong Gui, Yang FengNeurIPS 2024 · 19 citations
- EasySpec: Layer-Parallel Speculative Decoding for Efficient Multi-GPU UtilizationYize Wu, Ke Gao, Ling Li, Yanjun WuNeurIPS 2025 · 3 citations
- SPIN: Accelerating Large Language Model Inference with Heterogeneous Speculative ModelsFahao Chen, Peng Li, Tom H. Luan, Zhou Su et al.INFOCOM 2025 · 10 citations
- CoSine: Enhancing LLM Serving via Collaborative and Decoupled Speculative InferenceLuyao Gao, Jianchun Liu, Xichong Zhang, Guoju Gao et al.INFOCOM 2026 · 1 citation
