Strata: Hierarchical Context Caching for Long Context Language Model Serving
Zhiqiang Xie, Ziyi Xu, Mark Zhao, Yuwei An, Vikram Sharma Mailthody, Scott Mahlke, Michael Garland, Christos Kozyrakis
摘要
Long-context large language models (LLMs) enable applications that reason over hundreds of thousands to millions of tokens, but serving these workloads efficiently is challenging. Modern systems cache key-value (KV) states and rely on hierarchical context caching across GPU HBM, CPU memory, and SSDs. We show that naïve designs often become I/O-bound: fragmented KV layouts lead to small transfers that underutilize bandwidth, cache loading stalls prefill, and schedulers that ignore cache-loading latency and delay hits (concurrent requests for the same context during a cache miss) suffer severe throughput degradation. We present Strata, a hierarchical context caching framework for long-context LLM serving. Strata introduces a GPU-assisted I/O mechanism that decouples GPU and host layouts to enable efficient large transfers, and a cache-aware scheduler that mitigates delay hits, balances batches to hide cache-loading latency, and opportunistically overlaps complementary work. Implemented as part of SGLang and deployed in production, Strata improves throughput by up to 5× over vLLM-LMCache and 3.75× over NVIDIA TensorRT-LLM, without hurting short-context performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Agentic Plan Caching: Test-Time Memory for Fast and Cost-Efficient LLM AgentsQizheng Zhang, Michael Wornow, Kunle OlukotunNeurIPS 2025 · 被引用 27 次
- ThunderAgent: A Fast, Simple, and Program-Aware Agentic Inference SystemHao Kang, Ziyang Li, Xinyu Yang, Weili Xu 等ICML 2026 · 被引用 14 次
- KVDrive: A Holistic Multi-Tier KV Cache Management System for Long-Context LLM InferenceJian Lin, Jiazhi Mi, Zicong Hong, Haodong Wang 等SIGMOD 2026 · 被引用 5 次
- RepetitionCurse: Measuring and Understanding Router Imbalance in Mixture-of-Experts LLMs under DoS StressRuixuan Huang, Qingyue Wang, Hantao Huang, Yudong Gao 等ICML 2026
- ECHO: Efficient KV Cache Offloading with Lossless Prefetching for Serving Native Sparse Attention LLMsGuangda Liu, Wenhao Chen, Chengwei Li, Zhenyu Ning 等OSDI 2026
它引用的顶会 Paper21
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- SGLang: Efficient Execution of Structured Language Model ProgramsLianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun 等NeurIPS 2024 · 被引用 1,586 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- Orca: A Distributed Serving System for Transformer-Based Generative ModelsGyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim 等OSDI 2022 · 被引用 690 次
- DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model ServingYinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu 等OSDI 2024 · 被引用 646 次
相关 Paper
- ORBITFLOW: SLO-Aware Long-Context LLM Serving with Fine-Grained KV Cache ReconfigurationXinyue Ma, Heelim Hong, Taegeon Um, Jongseop Lee 等VLDB 2026 · 被引用 3 次
- S3: Increasing GPU Utilization during Generative Inference for Higher ThroughputYunho Jin, Chun-Feng Wu, David Brooks, Gu-Yeon WeiNeurIPS 2023 · 被引用 150 次
- Online Context Caching for Distributed Large Language Models ServingBin Gao, Zhuomin He, Yizhen Yao, Zhanzhi Lou Lou 等INFOCOM 2025 · 被引用 2 次
- ConServe: Fine-Grained GPU Harvesting for LLM Online and Offline Co-ServingYifan Qiao, Shan Yu, Shu Anzai, Haoran Ma 等ICML 2026 · 被引用 7 次
- SCBench: A KV Cache-Centric Analysis of Long-Context MethodsYucheng Li, Huiqiang Jiang, Qianhui Wu, Xufang Luo 等ICLR 2025
