Lune

EuroSys2025顶会

Stateful Large Language Model Serving with Pensieve

Lingfan Yu, Jinkun Lin, Jinyang Li

2025年份
23被引次数
28顶会引用

摘要

Large Language Models (LLMs) are wildly popular today and it is important to serve them efficiently. Existing LLM serving systems are stateless across requests. Consequently, when LLMs are used in the common setting of multi-turn conversations, a growing log of the conversation history must be processed alongside any request by the serving system at each turn, resulting in repeated processing.

In this paper, we design 𝑃𝑒𝑛𝑠𝑖𝑒𝑣𝑒, a system optimized for multi-turn conversation LLM serving. 𝑃𝑒𝑛𝑠𝑖𝑒𝑣𝑒 maintains the conversation state across requests by caching previously processed history to avoid duplicate processing. 𝑃𝑒𝑛𝑠𝑖𝑒𝑣𝑒's multi-tier caching strategy can utilize both GPU and CPU memory to efficiently store and retrieve cached data. 𝑃𝑒𝑛𝑠𝑖𝑒𝑣𝑒 also generalizes the recent PagedAttention kernel to support attention between multiple input tokens with a GPU cache spread over non-contiguous memory. Our evaluation shows that 𝑃𝑒𝑛𝑠𝑖𝑒𝑣𝑒 can achieve 1.14-3.0× the throughput of vLLM and TensorRT-LLM and significantly reduce latency.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper28

问问它们各自怎么用它

它引用的顶会 Paper22

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖