CoSine: Enhancing LLM Serving via Collaborative and Decoupled Speculative Inference
Luyao Gao, Jianchun Liu, Xichong Zhang, Guoju Gao, Yunming Liao
Abstract
Speculative inference accelerates Large Language Model (LLM) serving by leveraging small speculative models (SSMs) as drafters to generate draft tokens, which are subsequently verified in parallel by the LLM. However, the practical effectiveness of this approach is frequently hindered by the tight coupling of compute-intensive LLMs and memory-intensive SSMs within a single node (e.g., single-GPU device). This coupling induces resource contention, thereby significantly lowering draft acceptance rates and system throughput. To overcome these obstacles, we present CoSine, a novel speculative inference system that decouples sequential speculation from parallel verification. Specifically, CoSine distributes multiple specialized drafters across the heterogeneous nodes and coordinates them via a verification server, enabling cost-effective collaboration to mitigate inference bottlenecks. First, CoSine dynamically routes requests to the proper drafters and incorporates a confidence-based token fusion mechanism to synthesize high-quality drafts. Then, CoSine orchestrates the verification pipeline featuring prioritized draft queueing and adaptive batch management to balance the decoupled workflows and minimize pipeline bubbles. Experimental evaluation demonstrates that CoSine significantly outperforms state-of-the-art approaches. Under equivalent resource constraints, CoSine achieves latency reductions of up to 27.1% and throughput improvements from 1.24× to 1.62×.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 6743dfe5-bd83-4737-b194-0f257967fbe3Related papers
- SPIN: Accelerating Large Language Model Inference with Heterogeneous Speculative ModelsFahao Chen, Peng Li, Tom H. Luan, Zhou Su et al.INFOCOM 2025 · 10 citations
- SpecPIM: Accelerating Speculative Inference on PIM-Enabled System via Architecture-Dataflow Co-ExplorationCong Li, Zhe Zhou, Size Zheng, Jiaxi Zhang et al.ASPLOS 2024 · 29 citations
- GoodSpeed: Optimizing Fair Goodput with Adaptive Speculative Decoding in Distributed Edge InferencePhuong Tran, Tzu-Hao Liu, Long Tan Le, Tung-Anh Nguyen et al.INFOCOM 2026
- SpecInfer: Accelerating Large Language Model Serving with Tree-based Speculative Inference and VerificationXupeng Miao, Gabriele Oliaro, Zhihao Zhang, Xinhao Cheng et al.ASPLOS 2024 · 105 citations
- Towards Efficient LLM Inference via Collective and Adaptive Speculative DecodingSiqi Wang, Hailong Yang, Xuezhu Wang, Tongxuan Liu et al.SC 2025 · 3 citations
