CoSine: Enhancing LLM Serving via Collaborative and Decoupled Speculative Inference
Luyao Gao, Jianchun Liu, Xichong Zhang, Guoju Gao, Yunming Liao
摘要
Speculative inference accelerates Large Language Model (LLM) serving by leveraging small speculative models (SSMs) as drafters to generate draft tokens, which are subsequently verified in parallel by the LLM. However, the practical effectiveness of this approach is frequently hindered by the tight coupling of compute-intensive LLMs and memory-intensive SSMs within a single node (e.g., single-GPU device). This coupling induces resource contention, thereby significantly lowering draft acceptance rates and system throughput. To overcome these obstacles, we present CoSine, a novel speculative inference system that decouples sequential speculation from parallel verification. Specifically, CoSine distributes multiple specialized drafters across the heterogeneous nodes and coordinates them via a verification server, enabling cost-effective collaboration to mitigate inference bottlenecks. First, CoSine dynamically routes requests to the proper drafters and incorporates a confidence-based token fusion mechanism to synthesize high-quality drafts. Then, CoSine orchestrates the verification pipeline featuring prioritized draft queueing and adaptive batch management to balance the decoupled workflows and minimize pipeline bubbles. Experimental evaluation demonstrates that CoSine significantly outperforms state-of-the-art approaches. Under equivalent resource constraints, CoSine achieves latency reductions of up to 27.1% and throughput improvements from 1.24× to 1.62×.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- SPIN: Accelerating Large Language Model Inference with Heterogeneous Speculative ModelsFahao Chen, Peng Li, Tom H. Luan, Zhou Su 等INFOCOM 2025 · 被引用 10 次
- SpecPIM: Accelerating Speculative Inference on PIM-Enabled System via Architecture-Dataflow Co-ExplorationCong Li, Zhe Zhou, Size Zheng, Jiaxi Zhang 等ASPLOS 2024 · 被引用 29 次
- GoodSpeed: Optimizing Fair Goodput with Adaptive Speculative Decoding in Distributed Edge InferencePhuong Tran, Tzu-Hao Liu, Long Tan Le, Tung-Anh Nguyen 等INFOCOM 2026
- SpecInfer: Accelerating Large Language Model Serving with Tree-based Speculative Inference and VerificationXupeng Miao, Gabriele Oliaro, Zhihao Zhang, Xinhao Cheng 等ASPLOS 2024 · 被引用 105 次
- Towards Efficient LLM Inference via Collective and Adaptive Speculative DecodingSiqi Wang, Hailong Yang, Xuezhu Wang, Tongxuan Liu 等SC 2025 · 被引用 3 次
