Lune

EuroSys2025Top-tier venue

Stateful Large Language Model Serving with Pensieve

Lingfan Yu, Jinkun Lin, Jinyang Li

2025Year
23Citations
28Top-tier citations

Abstract

Large Language Models (LLMs) are wildly popular today and it is important to serve them efficiently. Existing LLM serving systems are stateless across requests. Consequently, when LLMs are used in the common setting of multi-turn conversations, a growing log of the conversation history must be processed alongside any request by the serving system at each turn, resulting in repeated processing.

In this paper, we design 𝑃𝑒𝑛𝑠𝑖𝑒𝑣𝑒, a system optimized for multi-turn conversation LLM serving. 𝑃𝑒𝑛𝑠𝑖𝑒𝑣𝑒 maintains the conversation state across requests by caching previously processed history to avoid duplicate processing. 𝑃𝑒𝑛𝑠𝑖𝑒𝑣𝑒's multi-tier caching strategy can utilize both GPU and CPU memory to efficiently store and retrieve cached data. 𝑃𝑒𝑛𝑠𝑖𝑒𝑣𝑒 also generalizes the recent PagedAttention kernel to support attention between multiple input tokens with a GPU cache spread over non-contiguous memory. Our evaluation shows that 𝑃𝑒𝑛𝑠𝑖𝑒𝑣𝑒 can achieve 1.14-3.0× the throughput of vLLM and TensorRT-LLM and significantly reduce latency.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

Cited by top-tier papers28

Ask how each one uses it

Builds on22

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines