Lune

INFOCOM2026顶会

CoSine: Enhancing LLM Serving via Collaborative and Decoupled Speculative Inference

Luyao Gao, Jianchun Liu, Xichong Zhang, Guoju Gao, Yunming Liao

2026年份
1被引次数

摘要

Speculative inference accelerates Large Language Model (LLM) serving by leveraging small speculative models (SSMs) as drafters to generate draft tokens, which are subsequently verified in parallel by the LLM. However, the practical effectiveness of this approach is frequently hindered by the tight coupling of compute-intensive LLMs and memory-intensive SSMs within a single node (e.g., single-GPU device). This coupling induces resource contention, thereby significantly lowering draft acceptance rates and system throughput. To overcome these obstacles, we present CoSine, a novel speculative inference system that decouples sequential speculation from parallel verification. Specifically, CoSine distributes multiple specialized drafters across the heterogeneous nodes and coordinates them via a verification server, enabling cost-effective collaboration to mitigate inference bottlenecks. First, CoSine dynamically routes requests to the proper drafters and incorporates a confidence-based token fusion mechanism to synthesize high-quality drafts. Then, CoSine orchestrates the verification pipeline featuring prioritized draft queueing and adaptive batch management to balance the decoupled workflows and minimize pipeline bubbles. Experimental evaluation demonstrates that CoSine significantly outperforms state-of-the-art approaches. Under equivalent resource constraints, CoSine achieves latency reductions of up to 27.1% and throughput improvements from 1.24× to 1.62×.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

lune papers get 6743dfe5-bd83-4737-b194-0f257967fbe3

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖